Llama
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
LLaMA-3 8B
1,020Decode Throughput (tok/s)
4
Llama-2 7B-Chat
76Latency (ms/token)
4
Llama 3.2 1B
1.94TPOTH
4
LLaMA linear layers (11008 × 4096) 7B
0.051Latency (ms)
4
LLaMA-7B linear layers (4096 × 11008)
0.051Latency (ms)
4
LLaMA-7B linear layers (4096 × 4096)
0.05Latency (ms)
4
LLaMA-13B (13824 × 5120 linear layer)
0.05Latency (ms)
4
LLaMA 5120 × 13824 linear layer 13B
0.051Latency (ms)
4
LLaMA-13B 5120 × 5120 linear layer
0.051Latency (ms)
4
LLaMA
0t-test
4
Llama 3.3 70B Emulation (train)
15.3Energy Reduction (Iso-Time)
4
Llama-8B
115.2Throughput (Tokens/s)
4
Llama-3B
215.6Throughput (TOK/s)
4
Llama-1B
310.5Throughput (Tokens/sec)
4
LLaMA 3
95ASR
4
LLaMA 2
100ASR
4
LLaMA2-13B
37Averaged Quantization Time (s)
4
Llama-2-7B 96k sequence length v1 (inference)
0.115Decoding Latency (s)
4
LLaMA offspring models
100MiGPT Score
4
Llama-3-8B
0.517Decode Time per Step
4
Llama 4
100Detection Accuracy (No Attack)
3
LLaMA 8B 3.1
6,517Sigma (ms)
3
LLaMA 7B
6,053Sigma Latency (ms)
3
Llama 8B 64K Context Length
2TTFT (s)
3
Llama 8B 32K Context Length
1TTFT (s)
3