Context Length
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Context-length sweep
6.008Full Perplexity (PPL)
8
32k context length efficiency Llama-3-8B (test)
4.12Time To First Token (s)
7
context length 128K
108Throughput (tok/s)
6
Context Length 32K
183Decode Throughput (tok/s)
6
8K Context Length
268.5Throughput (tok/s)
6
Context Length 32K
928Theoretical Compute (TFLOPs)
5
Context Length 16K
336Theoretical Compute (TFLOPs)
5
Context Length 4K
60Theoretical Compute (TFLOPs)
5
90k Context Length Llama-3.1-8B
8.9Throughput (queries/s)
4
30k Context Length (Llama-3.1-8B)
15.8Inference Throughput (QPS)
4
30k Context Length Llama-2-7B
6.6Inference Throughput (QPS)
4
Context Length 128K 1.0 (test)
2,360.6Effective b_KV (dense)
3
Context Length 32K 1.0 (test)
1,863.6Effective b_KV (dense)
3
Context Length 8K 1.0 (test)
1,658.3Effective KV Cache Size (dense)
3
Context Length 200K
10.7Prefill Time (s)
3
Context Length 120K
5.66Prefill Time (s)
3
Context Length 60K
2.59Prefill Time (s)
3
Context Length 10K
0.45Prefill Time (s)
3