Language Modeling
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
26.44PPL (0 Context)
3
Feb 26, 2026
20.72PPL (0 Context)
3
Feb 26, 2026
-36.23Relative P95 RTF Reduction
3
Feb 26, 2026
-30.21Relative P95 RTF Reduction
3
Feb 26, 2026
-2.11Relative P95 RTF Reduction
3
Feb 26, 2026
24.85Relative P95 RTF Reduction
3
Feb 26, 2026
3.41Relative P95 RTF Reduction
3
Feb 26, 2026
7.84Relative P95 RTF Reduction
3
Feb 26, 2026
4.66Relative P95 RTF Reduction
3
Feb 26, 2026
-23.79Relative P95 RTF Reduction
3
Feb 26, 2026
11.58PPL (Validation)
3
Feb 26, 2026
30.42PPL (All)
3
Feb 26, 2026
30.43PPL (All)
3
Feb 26, 2026
32.88Perplexity
3
Feb 26, 2026
41.94Perplexity
3
Feb 26, 2026
34.92Perplexity
3
Feb 26, 2026
40.92Perplexity
3
Feb 26, 2026
57.39Perplexity
3
Feb 26, 2026
2.48LM Loss
3
Feb 26, 2026
2.95Perplexity (PPL)
2
Jul 8, 2026
4.502Perplexity (PPL)
2
Jul 8, 2026
4.339Perplexity (PPL)
2
Jul 8, 2026
2.298Perplexity
2
Jul 8, 2026
2.304Perplexity
2
Jul 8, 2026
3.856Perplexity
2
Jul 8, 2026