Mathematical Reasoning on AIME 2024 (Avg@k and Pass@k Evaluation)
73.3Avg@1GPT-5 nano
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-5 nanoModel Version=nano2026.05 | 73.3 | 75 | 73.3 | 86.7 | |
| Qwen3-8BParameters=8B2026.05 | 70 | 71.2 | 70 | 86.7 | |
| gpt-oss-20bParameters=20b2026.05 | 70 | 67.5 | 70 | 83.3 | |
| Qwen3-4BParameters=4B2026.05 | 60 | 66.2 | 60 | 80 | |
| EngGPT2-16B-A3BParameters=16B2026.05 | 53.3 | 42.5 | 53.3 | 73.3 | |
| Ministral-3-8BParameters=8B2026.05 | 23.3 | 23.3 | 23.3 | 23.3 | |
| gemma-3-12b-itParameters=12b2026.05 | 16.7 | 18.3 | 16.7 | 23.3 | |
| gemma-3-4b-itParameters=4b2026.05 | 13.3 | 11.7 | 13.3 | 13.3 | |
| Llama-3.2-3B-InstructParameters=3B2026.05 | 3.3 | 3.3 | 3.3 | 3.3 | |
| Moonlight-16B-A3B-InstructParameters=16B2026.05 | 3.3 | 3.3 | 3.3 | 3.3 | |
| FastwebMIIA-7BParameters=7B2026.05 | 0 | 0 | 0 | 0 | |
| Minerva-7B-instruct-v1.0Parameters=7B2026.05 | 0 | 0 | 0 | 0 | |
| Velvet-14BParameters=14B2026.05 | 0 | 0 | 0 | 0 | |
| LLaMAntino-3-ANITA-8BParameters=8B2026.05 | 0 | 0 | 0 | 0 | |
| Llama-3.1-8B-InstructParameters=8B2026.05 | 0 | 0 | 0 | 0 | |
| deepseek-moe-16b-chatParameters=16b2026.05 | 0 | 0 | 0 | 0 |