Llama
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Llama-2-7B Taylor Pruning, 10% sparsity
0FSR
6
Llama3 TAR
32Success Rate First (SRF)
6
Llama3 RB
71Success Rate First (SRF)
6
Llama 3.1 8B activations
127.8Achieved L0
6
Llama Baseline 3.2-3B
78.65Accuracy
6
Llama-3.1-8B-Instruct Jailbreak Evaluation
0GCG Success Rate
6
Llama 3.1-8B-Instruct 512 -> 128 tokens, concurrency=1
269Throughput (Tok/s)
6
LLaMA 3B 3.2
7.81PPL
6
LLaMA 1B 3.2
9.75Perplexity (PPL)
6
LLaMA-3 8B
6.13PPL
6
LLaMA-2 13B
4.88Perplexity
6
LLaMA-2 7B
5.47PPL
6
Llama-3-8B Paraphrase perturbation, 150 tokens 1.0 (test)
0.041Mean P
6
Llama-3-8B Paraphrase perturbation, 30 tokens 1.0 (test)
0.21Mean P
6
Llama-3-8B Translate perturbation, 150 tokens 1.0 (test)
0.093Mean P
6
Llama-3-8B Delete perturbation, 150 tokens 1.0 (test)
0.29Mean P
6
Llama-3-8B Swap perturbation, 150 tokens 1.0 (test)
0.01Mean P
6
Llama-3-8B Delete 50%, 150 Tokens
0.36Mean P
6
Llama-3-8B Delete 50%, 30 Tokens
0.38Mean P
6
Llama-3-8B Delete 30%, 150 Tokens
0.31Mean P
6
Llama-3-8B Delete 30%, 30 Tokens
0.26Mean P
6
Llama-3-8B Swap 50%, 150 Tokens
0.12Mean P
6
Llama-3-8B Swap 30%, 150 Tokens
0.03Mean P
6
Llama-3-8B Swap 30%, 30 Tokens
0.16Mean P
6
Llama (test)
14.77Throughput (tokens/s)
6