Long-Context
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Long Context 32K
1,218.1Decode Throughput (tok/s)
48
Long Context 64K
1,042.6Decoding Throughput (tok/s)
42
Long Context 96K
930.5Decode Throughput (tok/s)
38
Long Context 128K
876.9Throughput (tok/s)
33
Long-context benchmarks
74.2ICL Performance (8k Context)
21
Long-context (test)
1mTokens
19
Long Context benchmark
67.59Accuracy
14
Long-context 15 datasets v2 (test)
0.523Avg. Normalized RMSE
9
Long-context benchmarks
45.9Performance (8k Context)
8
Long-context 1024-token input, 32-token output
1.48TPOT Speedup vs DeepGEMM
3
Long-Context (train)
—Primary metric
0