Mathematical Reasoning on AIME 2025 (Avg@1/8 e_m, Pass@1/8 e_m)
72.9Avg@8 (e_m)GPT-5 nano
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-5 nano2026.05 | 72.9 | 73.3 | 73.3 | 86.7 | |
| gpt-oss-20b2026.05 | 66.7 | 76.7 | 76.7 | 93.3 | |
| Qwen3-8B2026.05 | 61.2 | 63.3 | 63.3 | 76.7 | |
| Qwen3-4B2026.05 | 58.3 | 60 | 60 | 70 | |
| EngGPT2-16B-A3B2026.05 | 31.7 | 30 | 30 | 50 | |
| Ministral-3-8B2026.05 | 26.7 | 26.7 | 26.7 | 26.7 | |
| GRPOBase Model=Qwen3-4B2026.05 | 25.8 | — | — | 36.66 | |
| MOPDBase Model=Qwen3-4B2026.05 | 25.41 | — | — | 36.28 | |
| Qwen3-4Bstatus=base model2026.05 | 17.92 | — | — | 32.57 | |
| gemma-3-12b-it2026.05 | 17.5 | 13.3 | 13.3 | 20 | |
| gemma-3-4b-it2026.05 | 13.3 | 13.3 | 13.3 | 13.3 | |
| Moonlight-16B-A3B-Instruct2026.05 | 9.6 | 6.7 | 6.7 | 10 | |
| SDPOBase Model=Qwen3-4B2026.05 | 7.81 | — | — | 16.66 | |
| FastwebMIIA-7B2026.05 | 0 | 0 | 0 | 0 | |
| Minerva-7B-instruct-v1.02026.05 | 0 | 0 | 0 | 0 | |
| Velvet-14B2026.05 | 0 | 0 | 0 | 0 | |
| LLaMAntino-3-ANITA-8B2026.05 | 0 | 0 | 0 | 0 | |
| Llama-3.2-3B-Instruct2026.05 | 0 | 0 | 0 | 0 | |
| Llama-3.1-8B-Instruct2026.05 | 0 | 0 | 0 | 0 | |
| deepseek-moe-16b-chat2026.05 | 0 | 0 | 0 | 0 |