Question Answering on SimpleQA (Acc, BAS, ECE, AURC)
32.9AccuracyGrok-3
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Grok-3Model scale group=Frontier / Large Models (> 100B params)2026.04 | 32.9 | -0.9 | 60.1 | 65 | |
| DS-R1Model scale group=Frontier / Large Models (> 100B params)2026.04 | 29.9 | -0.48 | 52.7 | 60 | |
| Mist(L)Model scale group=Frontier / Large Models (> 100B params)2026.04 | 28.7 | -1.09 | 64.2 | 68 | |
| GPT-4oModel scale group=Frontier / Large Models (> 100B params)2026.04 | 21.4 | -1.3 | 70.5 | 76 | |
| Mist(M)Model scale group=Mid-tier Models (10B to 70B params)2026.04 | 20.2 | -1.26 | 68.8 | 74 | |
| G3 MiniModel scale group=Mid-tier Models (10B to 70B params)2026.04 | 19.6 | -0.86 | 60.8 | 71 | |
| Llama 3.3Model scale group=Mid-tier Models (10B to 70B params)2026.04 | 19.4 | -2.97 | 68.1 | 75 | |
| DS-V3.2Model scale group=Frontier / Large Models (> 100B params)2026.04 | 17.6 | -0.63 | 49 | 75 | |
| GPT-o1Model scale group=Frontier / Large Models (> 100B params)2026.04 | 12 | -0.66 | 58.5 | 77 | |
| Mist(S)Model scale group=Small / Efficient Models (<10B params or highly optimized)2026.04 | 10.9 | -1.63 | 82.6 | 89 | |
| 4o-miniModel scale group=Small / Efficient Models (<10B params or highly optimized)2026.04 | 8.9 | -3.86 | 84.2 | 89 | |
| Phi-4Model scale group=Small / Efficient Models (<10B params or highly optimized)2026.04 | 8.5 | -3.63 | 84.1 | 91 | |
| Ours (+GDPO)Alignment=GDPO2026.05 | 6.45 | — | — | — | |
| STAIR2026.05 | 6.38 | — | — | — | |
| Shallow-Align2026.05 | 5.2 | — | — | — | |
| DPO2026.05 | 4.46 | — | — | — | |
| SFT2026.05 | 4.27 | — | — | — | |
| Ours (+SFT)Alignment=SFT2026.05 | 4.21 | — | — | — | |
| Self-Critique2026.05 | 4.09 | — | — | — | |
| Original2026.05 | 2.52 | — | — | — |