LongBench
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
LongBench V2 (test)
60Acc (Short)
7
LongBench Llama-3.2-1B-Instruct (test)
16.12NQA
7
LongBench MuSiQue and WikiMultiHopQA
69.9F1 Score
7
LongBench-Write-en [4k, 20k)
21.5Sl Score
7
LongBench-Write-en [0, 500)
92.1Sl Score
7
LongBench
50.63Average Score
6
LongBench Code Repo QA Long v2
72.41Accuracy
6
LongBench 16K context
19.02LongBench Score (16k)
6
LongBench
96.93Throughput (RPS)
6
LongBench 16K context length
13.3NrtvQA Score
6
LongBench (standard)
5.89NQA
6
LongBench 2
50.3P@4
6
LongBench single-node
0.79TTFT (Mean)
6
LongBench 16K length
50.5LCC
6
LongBench
87.3LB Score
6
LongBench
30.4NarQA Score
6
LongBench V1
45.77Qasper
6
LongBench
62.5SR (Stability)
6
LongBench-Zh filtered (test)
64.8Single Document Score
5
LongBench Zh
62.24Single Document Performance
5
LongBench En
38.84Single-Document Score
5
LongBench
25.19Qasper Score
5
LongBench Avg.
43.99LongBench Avg Score
5
LongBench Code Completion
59.78Accuracy
5
LongBench Synthetic Tasks
67.78Accuracy
5