Mathematical Reasoning on AIME 2025 (Pass@1, #Tokens)
93.3Pass@1 AccuracyGPT-5(high)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5(high)Model Scale=Reference, Model Category=Closed-source frontier2026.06 | 93.3 | — | |
| DeepSeek-V4-proModel Scale=Reference, Model Category=Open-source frontier2026.06 | 93.3 | — | |
| DeepSeek-R1Model Scale=Reference, Model Category=Open-source frontier2026.06 | 90 | — | |
| Qwen3-maxModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 80 | — | |
| REBALANCEHardware=Ascend 910B NPU, Backbone=openPangu-Embedded-7B-V1.1, Thinking Mode=slow-thinking mode2026.03 | 76.7 | 9,417 | |
| BaselineHardware=Ascend 910B NPU, Backbone=openPangu-Embedded-7B-V1.1, Thinking Mode=slow-thinking mode2026.03 | 73.3 | 14,552 | |
| Qwen3-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 73.3 | — | |
| DiScOw/o ITDEModel Scale=32B, Model Category=Ours2026.06 | 66.7 | — | |
| DiScOModel Scale=32B, Model Category=Ours2026.06 | 66.7 | — | |
| DeepSeek-R1-Distill-Qwen-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 60 | — | |
| TOPDBase Model=Qwen3-4B (Warm Start)2026.05 | 53.3 | — | |
| GPT-o1-miniModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 50.8 | — | |
| DiScOw/o ITDEModel Scale=7B, Model Category=Ours2026.06 | 50 | — | |
| DiScOModel Scale=7B, Model Category=Ours2026.06 | 50 | — | |
| OPDBase Model=Qwen3-4B (Warm Start)2026.05 | 46.7 | — | |
| GPT-o1-previewModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 46.7 | — | |
| Qwen3-8BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 46.7 | — | |
| rStar-Math-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 46.7 | — | |
| QwQ-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 46.7 | — | |
| ICPOBase Model=8B, Training Method=ICPO2025.10 | 43.7 | — | |
| Claude3.5-SonnetModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 43.3 | — | |
| ICPO†Base Model=8B, Training Method=ICPO†2025.10 | 40.9 | — | |
| GRPOExpertDomainBase Model=8B, Training Method=GRPOExpertDomain2025.10 | 40.4 | — | |
| GRPOExtraRolloutsBase Model=8B, Training Method=GRPOExtraRollouts2025.10 | 40 | — | |
| Qwen3-4BStage=Warm Start2026.05 | 40 | — | |
| DeepSeek-V3Model Scale=Reference, Model Category=Open-source frontier2026.06 | 39.2 | — | |
| GRPOBase Model=8B, Training Method=GRPO2025.10 | 38.5 | — | |
| ReasonFlux-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 37.2 | — | |
| ReasonFlux-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 36.7 | — | |
| Sky-T1-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 36.7 | — | |
| DeepSearch-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 35.42 | — | |
| Nemotron-Research-Reasoning-Qwen-1.5B v1number of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 33.85 | — | |
| Nemotron-Research-Reasoning-Qwen-1.5B v2number of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 32.92 | — | |
| DeepScaleR-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 30.52 | — | |
| LoRA-αRank (r)=256, n=642026.06 | 28.95 | — | |
| Full Fine-TuningRank (r)=N/A, n=642026.06 | 28.69 | — | |
| LoRA (10× LR)Rank (r)=256, n=642026.06 | 26.71 | — | |
| GPT-4oModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 26.7 | — | |
| DeepSeek-Coder-V2-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 26.7 | — | |
| Qwen2.5-Math-72B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 26.7 | — | |
| ICPO†Base Model=1.7B, Training Method=ICPO†2025.10 | 26.6 | — | |
| ICPOBase Model=1.7B, Training Method=ICPO2025.10 | 26.3 | — | |
| LoRA (10× LR)Rank (r)=64, n=642026.06 | 25.83 | — | |
| LoRA-αRank (r)=64, n=642026.06 | 25.83 | — | |
| STILL-3-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 25 | — | |
| GRPOExpertDomainBase Model=1.7B, Training Method=GRPOExpertDomain2025.10 | 24.8 | — | |
| Open-RS3-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 24.79 | — | |
| GRPOExtraRolloutsBase Model=1.7B, Training Method=GRPOExtraRollouts2025.10 | 24.7 | — | |
| Open-RS2-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 24.37 | — | |
| DeepSeek-R1-Distill-Qwen-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 24.06 | — | |
| LoRARank (r)=256, n=642026.06 | 23.69 | — | |
| Open-RS1-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 22.6 | — | |
| GRPOBase Model=1.7B, Training Method=GRPO2025.10 | 22.5 | — | |
| S2L-POFamily=Qwen3, Configuration=1.7B→8B2026.05 | 22.5 | — | |
| DIVERModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 22.5 | — | |
| Qwen3Base Model=8B, Training Method=Base2025.10 | 21.4 | — | |
| Qwen3Base Model=1.7B, Training Method=Base2025.10 | 20.6 | — | |
| LoRARank (r)=64, n=642026.06 | 19.73 | — | |
| Qwen2.5-Math-7B-InstructModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 16.7 | — | |
| GRPO w/ Clip-higherModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 16.7 | — | |
| Pass@k TrainingModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 16.6 | — | |
| S2L-POFamily=Qwen3, Configuration=4B→14B2026.05 | 14.6 | — | |
| NuminaMath-72B-CoTModel Scale=Reference, Model Category=Open-source frontier2026.06 | 13.3 | — | |
| SuperCorrect-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 13.3 | — | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 13.3 | — | |
| Entropy-RLModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 13.3 | — | |
| Qwen2.5-32B-InstructModel Scale=32B, Model Category=Instruction-tuned models2026.06 | 13.3 | — | |
| GRPOFamily=Qwen3, Model Size=14B2026.05 | 12.9 | — | |
| GRPOFamily=Qwen3, Model Size=8B2026.05 | 12.1 | — | |
| Qwen2.5-Math-1.5B-Oat-Zeronumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 10 | — | |
| LLaMA3.1-405B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 10 | — | |
| Qwen2.5-Math-1.5B-Instructnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 8.85 | — | |
| Qwen2.5-Math-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 6.35 | — | |
| S2L-POFamily=InternLM2.5, Configuration=1.8B→7B2026.05 | 3.5 | — | |
| LLaMA3.1-70B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 3.3 | — | |
| Qwen2.5-Math-7BModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 3.3 | — | |
| GRPOFamily=InternLM2.5, Model Size=7B2026.05 | 0.1 | — | |
| LLaMA3.1-8B-InstructModel Scale=7B, Model Category=General instruction-tuned models2026.06 | 0 | — | |
| Mathstral-7B-v0.1Model Scale=7B, Model Category=General instruction-tuned models2026.06 | 0 | — |