General Language Model Evaluation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
68.93MMLU
34
Apr 16, 2026
74.9Average Accuracy
21
May 21, 2026
50.7Overall Average Score
14
Jun 2, 2026
45.02Overall Performance
12
Feb 26, 2026
33.7LBPP Score
10
Apr 7, 2026
7.03Win Rate
9
Feb 26, 2026
26.95WildBench Score
2
Feb 26, 2026