Llama 2
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Llama-2-7b-chat-hf 10 samples UMD watermarking (test)
1AUROC (t=0)
64
LLaMA-2 Chat 7B
0.075Attention Latency (ms)
60
LLaMA-2-7B-CHAT Safety (test)
0.55Safety Score
60
Llama-2 13B
4.85Perplexity (PPL)
32
Llama-2-7b-chat finetuned variants v1 (test)
60.4Transfer Success Rate (TSR)
16
Llama-2-7b-chat-hf UMD watermarking (10 samples)
100ASR
15
Llama-2-70B
2.2GPU Hours (h)
13
Llama-2 7B
0.17Verification Latency (s)
12
Llama-2 7B pre-training
0Number of Spikes
9
LLaMA-2 32B
0.253Reconfiguration Time (s)
8
Llama-2-7B teacher vs. llama-2-7b-logit-watermark-distill-kgw-k1-gamma0.25-delta2 student (test)
99.98Similarity Score
7
Llama-2 DPO 7B
99.94Similarity Score
7
Llama-2-7b-Chat-hf Open-Ended Generation
2.46Wealth Score
7
Llama-2 70B
36.1Throughput (tokens/s)
6
Llama-2-7B SFT & RLHF
2FSR (Anchor)
6
Llama-2-7B Taylor Pruning 5% sparsity
0FSR
6
Llama-2-7B Random Pruning, 10% sparsity
0FSR
6
Llama-2-7B Random Pruning, 5% sparsity
0False Success Rate (FSR)
6
Llama-2-7B 32k sequence length v1 (inference)
0.062Decoding Latency (s)
6
Llama-2-7B 16k sequence length v1 (inference)
0.041Decoding Latency (s)
6
Llama 2 7B inference v1.0
188Decoding Throughput (TOK/s)
6
LLaMA-2 7B pre-training (val)
16.01Validation Perplexity (40K steps)
5
LLaMA-2 7B fine-tuned variants
0U-test p-value
5
LLaMA-2 70B sequence length 2048
384Max Batch Size
5
Llama-2-7B 64k sequence length v1 (inference)
0.098Decoding Latency (s)
5