Average
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Average Protocol I
59.78IoU
9
Average (test)
44.84BLEU-1
9
Average All Benchmarks
71Accuracy
9
Average (HumanEval, GSM8K, Math-500, MT-Bench)
4.82SpeedUp
9
Average (Alpaca, Samsum, ChatDoctor)
0.508BRT Score
9
Average MuSiQue, 2WikiMQA, MFQA-En, NarrativeQA, En.QA
42.48SubEM Score
9
Average 11 datasets (test)
85.74Base Accuracy
9
Average Boulder, Ottawa, Hawaii
66.23phi_EN
9
Average StrategyQA, MMLU, TruthfulQA, ARC-Challenge
80.2AUROC
9
Average MATH500, AMC, AIME24, GPQA
43.8Acc @ First
9
Average Public Datasets v1 (test)
24.74PSNR
9
Average 494
27.21PSNR
9
Average
0.66Dice Coefficient
9
Average Pythia (train)
51.4TPR@1%FPR
9
Average Aggregate of AIME24, AIME25, AMC23, MATH500, OlympiadBench
61.03Pass@1
9
Average
94.7Avg Performance
9
Average AIME24, AIME25, MATH500, AMC23, Hmmt25, Olympiad
30.07Average Performance
8
Average
76.6Human-Model Similarity (1-nWD)
8
Average AIME 2025, BeyondAIME, HMMT 02/25, HMMT 02/26, MATH-500
26.7Pass@1
8
Average
54.1Avg @1 Score
8
Average Across Simulation and Real-world
53.6Success Rate
8
Average Nature, Real, SIR2
27.24PSNR
8
Average
33.04Avg PD
8
Average
16.25BLEU
8
Average 11 FAS datasets
6.3HTER (%)
8