Llama
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Llama 405B (128 Q-heads/8 KV-heads/128 Head-dimension) 3.1
615.39TFLOPS
62
LLaMA-2-7B-Chat
449Throughput (tokens/sec)
60
Llama S=512 on B200 3.1-8B (base variant)
3.1Peak Memory Usage (GiB)
58
Llama-3.2-1B-Instruct Task Set
0.83Score S(gamma)
56
Llama 8B base variant on H200 3.1
99.1Per-step Latency (ms)
49
LLaMA-8B-Instruct Chunked Prefill 3.1 (inference)
423.1Attention Latency (ms)
49
Llama 70B 3.1
3,119.55Throughput
48
Llama-3-8B n≈200
77ASR
42
Llama-3-8B decoder block
153Latency (µs)
36
LLAMA 1
88Accuracy
34
LLaMA-2 7B (inference)
8.23P99 Per-Token Latency (ms)
33
LLAMA
0.25Processing Time (hr)
30
Llama 3.1 8B Q4 weights
204TTFT (ms)
28
Llama 70B (H100 GPU Cluster) 3.1
894.32Throughput
27
Llama2 7B
1Similarity Score
24
Llama-8B model tree (test)
1Rank
21
Llama 70B 3.1 (inference)
1,410.39Throughput
21
Llama 8B 3.1
96NR Rate
20
LLaMA-2 13B
19.4Throughput (tokens/s)
20
LLaMA 13B 2
4.57Perplexity (PPL)
20
Llama-3.1-70B Large Target
98Similarity Score
18
Llama 8B Small Target 3.1
0.97Similarity Score
18
Llama2-7B evaluation scenarios (test)
85.16Accuracy
18
LLaMA-2-7B
5.47Perplexity
18
Llama-3B Target Transferability set
81ASR
17