Competitive Mathematical Reasoning on AIME 2024 (Pass@1)
100Pass@1 AccuracyDeepSeek-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1Model Scale=Reference, Model Category=Open-source frontier2026.06 | 100 | |
| GPT-5(high)Model Scale=Reference, Model Category=Closed-source frontier2026.06 | 96.7 | |
| DeepSeek-V4-proModel Scale=Reference, Model Category=Open-source frontier2026.06 | 96.7 | |
| Qwen3-maxModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 93.3 | |
| DiScOw/o ITDEModel Scale=32B, Model Category=Ours2026.06 | 86.7 | |
| DiScOModel Scale=32B, Model Category=Ours2026.06 | 86.7 | |
| Qwen3-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 83.3 | |
| DeepSeek-R1-Distill-Qwen-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 80 | |
| DiScOModel Scale=7B, Model Category=Ours2026.06 | 70 | |
| DiScOw/o ITDEModel Scale=7B, Model Category=Ours2026.06 | 66.7 | |
| GPT-o1-miniModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 56.7 | |
| Qwen3-8BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 56.7 | |
| ReasonFlux-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 56.7 | |
| rStar-Math-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 50 | |
| QwQ-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 50 | |
| DeepSeek-V3Model Scale=Reference, Model Category=Open-source frontier2026.06 | 47.7 | |
| GPT-o1-previewModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 44.6 | |
| Sky-T1-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 43.3 | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 36.7 | |
| ReasonFlux-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 36.7 | |
| Qwen2.5-Math-72B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 30 | |
| SuperCorrect-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 26.7 | |
| DIVERModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 25 | |
| LLaMA3.1-405B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 23.3 | |
| Entropy-RLModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 23.3 | |
| GRPO w/ Clip-higherModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 20 | |
| Pass@k TrainingModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 20 | |
| GPT-4oModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 16.7 | |
| LLaMA3.1-70B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 16.7 | |
| Qwen2.5-Math-7BModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 16.7 | |
| Qwen2.5-32B-InstructModel Scale=32B, Model Category=Instruction-tuned models2026.06 | 16.5 | |
| Claude3.5-SonnetModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 16 | |
| DeepSeek-Coder-V2-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 13.3 | |
| Qwen2.5-Math-7B-InstructModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 13.3 | |
| LLaMA3.1-8B-InstructModel Scale=7B, Model Category=General instruction-tuned models2026.06 | 6.7 | |
| NuminaMath-72B-CoTModel Scale=Reference, Model Category=Open-source frontier2026.06 | 3.3 | |
| Mathstral-7B-v0.1Model Scale=7B, Model Category=General instruction-tuned models2026.06 | 0 |