Long-context Reasoning
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
99.25NIAH
6
May 20, 2026
50.3P@4
6
May 12, 2026
48.5Accuracy
6
Feb 26, 2026
74.7Score
5
Jun 23, 2026
89.77Oolong Score
5
Jun 12, 2026
78MTP Acceptance Rate
5
Jun 11, 2026
66.9Accuracy
5
Apr 29, 2026
29.33MultiFieldQA Score
5
Feb 26, 2026
0.32Average Reward
4
May 8, 2026
46.2QA Score
4
Mar 24, 2026
54.64Average LC Score
3
Jun 18, 2026
92.78NIAH
3
May 20, 2026
91.64Accuracy
3
Apr 15, 2026
95.95Accuracy (RULER 512k)
3
Apr 15, 2026
96.83Accuracy
3
Apr 15, 2026
100Single-NIAH Score
3
Apr 14, 2026
100Single-NIAH Score
3
Apr 14, 2026
100Single-NIAH
3
Apr 14, 2026
0.249Average Reward
2
May 8, 2026
45.4Average Reward
2
May 8, 2026