Competitive Mathematical Reasoning on AMC 2023 (Pass@1)
100Pass@1 AccuracyQwen3-max
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-maxModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 100 | |
| DeepSeek-V4-proModel Scale=Reference, Model Category=Open-source frontier2026.06 | 100 | |
| GPT-5(high)Model Scale=Reference, Model Category=Closed-source frontier2026.06 | 97.5 | |
| DeepSeek-R1-Distill-Qwen-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 97.5 | |
| DiScOw/o ITDEModel Scale=32B, Model Category=Ours2026.06 | 97.5 | |
| DiScOModel Scale=32B, Model Category=Ours2026.06 | 97.5 | |
| DiScOModel Scale=7B, Model Category=Ours2026.06 | 92.5 | |
| Qwen3-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 92.5 | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 87.5 | |
| rStar-Math-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 87.5 | |
| DiScOw/o ITDEModel Scale=7B, Model Category=Ours2026.06 | 87.5 | |
| GPT-o1-miniModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 85 | |
| GPT-o1-previewModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 85 | |
| ReasonFlux-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 85 | |
| Sky-T1-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 82.5 | |
| Claude3.5-SonnetModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 80 | |
| DeepSeek-V3Model Scale=Reference, Model Category=Open-source frontier2026.06 | 80 | |
| ReasonFlux-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 80 | |
| QwQ-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 75 | |
| NuminaMath-72B-CoTModel Scale=Reference, Model Category=Open-source frontier2026.06 | 70 | |
| Qwen2.5-Math-72B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 70 | |
| Qwen3-8BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 65 | |
| DeepSeek-R1Model Scale=Reference, Model Category=Open-source frontier2026.06 | 64.1 | |
| Qwen2.5-32B-InstructModel Scale=32B, Model Category=Instruction-tuned models2026.06 | 64 | |
| DIVERModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 62.5 | |
| DeepSeek-Coder-V2-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 57.5 | |
| Entropy-RLModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 57.5 | |
| GRPO w/ Clip-higherModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 57.5 | |
| Pass@k TrainingModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 52.5 | |
| LLaMA3.1-70B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 50 | |
| LLaMA3.1-405B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 50 | |
| GPT-4oModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 47.5 | |
| Mathstral-7B-v0.1Model Scale=7B, Model Category=General instruction-tuned models2026.06 | 37.5 | |
| SuperCorrect-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 37.5 | |
| Qwen2.5-Math-7B-InstructModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 28.3 | |
| LLaMA3.1-8B-InstructModel Scale=7B, Model Category=General instruction-tuned models2026.06 | 25 | |
| Qwen2.5-Math-7BModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 22.5 |