Llama
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Style poisoned Llama 3.1 8B Instruct
5.44Utility Score
7
Llama-2 7B base
4.753Perplexity
7
LLaMA 8B Instruct 128K context length 3.1
24,055.4Attention Latency (ms)
7
LLaMA 8B-Instruct 64K context length 3.1
8,945.9Attention Latency
7
LLaMA 8B Instruct 16K context length 3.1
1,224Attention Prefill Latency
7
Llama-3.1-8B teacher vs. Llama-3.1-8B-Instruct-Open-R1-Distill student (test)
99.88Similarity Score
7
Llama-2-7B ppo-v0.1-reward
100Similarity Score
7
Llama3.1-8B (train)
20.19Memory (GB)
7
LLaMA-8B-J unseen instances
0.756Hypervolume
7
LLaMA-8B-G unseen instances
79Hypervolume
7
LLaMA-7B-J unseen instances
0.709Hypervolume
7
LLaMA-7B-G (unseen instances)
76.2Hypervolume
7
Llama 3.1 8B (val)
0.001DKL
7
Llama 8B 3.1
19Compression Time
7
Llama 8B 3.1
2.8Model Size (GB)
7
Llama 3.2 3B
5.52Avg VRAM (MB)
7
Llama-3.1-8B 32k sequence length v1 (inference)
0.033Decoding Latency (s)
7
Llama2
215ATGR
7
LLaMA2-7B-Chat
1,800Detection Percentage
7
LLaMA-2 Token Replacement Attack, epsilon=0.2 (1,000 generated sequences)
88.79TPR @ FPR=0.1%
7
LLaMA-2 Token Replacement Attack epsilon=0.1 (1,000 generated sequences)
96.07TPR@FPR=0.1%
7
LLaMA-2 Token Replacement Attack epsilon=0.05 (1,000 generated sequences)
97.11TPR@FPR=0.1%
7
Llama 3B 3.2
117.3Throughput (tok/s)
6
Llama 3.2 1B
295.4Throughput (TOK/s)
6
Llama 3.1 8B
99.281DIPMark
6