Llama
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
LLaMA Chat 1B
0.001vNMSE
6
LLaMA3 8B Instruct
88GCG Success Rate
6
LLaMA-2 7B HF
25.4S@>=1
6
Llama
100Detectability (Origin)
5
Llama-3-8B
130.7Throughput (Tok/s)
5
Llama 3.1-8b SAE llamascope-res-32k Layer 2
52.33Gen Score
5
Llama3.1-8b SAE llamascope-res-32k Layer 1
58.4GEN
5
Llama-3.3-70B math n=30 (failure pool)
86.7Correction Rate
5
LLaMA-3-1B Zero-shot
9.6Perplexity (PPL)
5
LLaMA 1B pre-training 2 (val)
15.22Perplexity
5
Llama 3.1 8B
225.01Throughput (TOK/s)
5
LLaMA 8B 100 clients 3.1
4.83Perplexity (PPL)
5
Llama-3.2-11B-Vision-Instruct (test)
41.85Detoxify Score
5
Llama-3-8B
63.04Coverage
5
Llama 7B 3.1
64.25Benign F1 Score
5
LLaMA 8B 3.1
150.8Token Throughput (4K Context) (tokens/sec)
5
Llama-2 7B
62.9HV Score
5
Llama-3-8B
4.22Model Size (GB)
5
LLaMA-2-7B
7.42Weights Memory Usage
5
LLaMA 8B 32K context length 3.1
1,115Theoretical Compute (TFLOPs)
5
Llama 3.1 8B (2048 samples)
62.8Training Time (s)
5
Llama (held-out test)
1.21Throughput Ratio Improvement
5
LLaMA 2-13B (inference)
96.02Attack Detection Accuracy
5
Llama-3.1-8B
7.3TPOT
5
Llama 4 Maverick
46ASR
5