Long-context language modeling
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
13.3NrtvQA Score
6
May 20, 2026
79.39Quality Score
6
May 20, 2026
10.21CWE (Context Length 4K)
6
May 12, 2026
3.93NLL (Context=1024)
5
May 25, 2026
78.7Average Score
5
Mar 13, 2026
83.81Accuracy (16K Context)
4
May 20, 2026
4.14Book Perplexity
4
Feb 26, 2026
80.12Q Score (Dense)
3
May 20, 2026
79.39Q* Score (Dense)
3
May 20, 2026
81.35Q* Score (Dense)
3
May 20, 2026
100S-NIAH Component 1 Score
2
May 12, 2026
100S-NIAH Component 1
2
May 12, 2026
81.5Accuracy (8k Context)
2
Feb 26, 2026