Mathematical Reasoning on AIME (held-out) (accuracy)
66.7AccuracyMAPPA
Evaluation Results
| Method | Links | |
|---|---|---|
| MAPPABackbone=Qwen3-4B, Model Setup=Best2026.01 | 66.7 | |
| TRINITYZero-shot=true2025.12 | 50 | |
| Qwen3-4BModel Setup=Baseline2026.01 | 49.2 | |
| Gemini Pro 2.5Zero-shot=true2025.12 | 46.67 | |
| GPT-5Zero-shot=true2025.12 | 46.67 | |
| Claude-4-SonnetZero-shot=true2025.12 | 35.33 | |
| DeepSeek-R1-Qwen-32BZero-shot=true2025.12 | 30 | |
| MAPPABackbone=R1-Distill-Qwen-1.5B, Model Setup=Best2026.01 | 29.2 | |
| R1-Distill-Qwen-1.5BModel Setup=Baseline2026.01 | 24.2 | |
| Qwen3-32BZero-shot=true, Evaluation Mode=reasoning2025.12 | 23.33 | |
| Qwen3-32BZero-shot=true, Evaluation Mode=direct2025.12 | 20 | |
| Gemma-3-27B-ITZero-shot=true2025.12 | 20 |