Mathematical Reasoning on AIME 24 (Pass@1 Accuracy, Avg Output Length)
76.25Pass@1 AccuracyQwen3-30B-A3B (teacher)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-30B-A3B (teacher)Access=–2026.05 | 76.25 | — | |
| ROPDAccess=text2026.05 | 63.33 | — | |
| ESampModel=GPT-OSS-20B2026.04 | 62.7 | — | |
| OverRIDEModel=GPT-OSS-20B2026.04 | 58.8 | — | |
| Min-pModel=GPT-OSS-20B2026.04 | 58.7 | — | |
| VanillaBackbone=Qwen3-4B2026.03 | 57.9 | 11,933 | |
| FIREModel=GPT-OSS-20B2026.04 | 57.8 | — | |
| VanillaModel=GPT-OSS-20B2026.04 | 57.2 | — | |
| TRiMSBackbone=Qwen3-4B2026.03 | 53.5 | 1,713.6 | |
| HölderPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 53.3 | — | |
| ExOPDAccess=logit2026.05 | 50.66 | — | |
| Dr.GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 50 | — | |
| NUDGERLModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 48.2 | — | |
| LOPDAccess=logit2026.05 | 47.92 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=3, p-Scheduling Strategy=static2026.05 | 46.7 | — | |
| GMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=p → 02026.05 | 46.7 | — | |
| PMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 46.7 | — | |
| POPEModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 46 | — | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=162026.05 | 45.4 | — | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=322026.05 | 45.1 | — | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 44.4 | — | |
| Oat-Zero-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 43.3 | — | |
| GMPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=p → 02026.05 | 43.3 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=p → 0, p-Scheduling Strategy=dynamic2026.05 | 43.3 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=2, p-Scheduling Strategy=static2026.05 | 43.3 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 43.3 | — | |
| GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=12026.05 | 43.3 | — | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=642026.05 | 41.5 | — | |
| ESampModel=Qwen3-8B2026.04 | 41 | — | |
| FIREModel=Qwen3-8B2026.04 | 40.4 | — | |
| Still-3-1.5B-PreviewZero-shot=true, Model Category=Qwen 1.5B Models, Training Samples=30K2025.05 | 40 | — | |
| GRPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=12026.05 | 40 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=-1, p-Scheduling Strategy=static2026.05 | 40 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=1, p-Scheduling Strategy=static2026.05 | 40 | — | |
| Min-pModel=Qwen3-8B2026.04 | 38.4 | — | |
| VanillaModel=Qwen3-8B2026.04 | 38.1 | — | |
| OverRIDEModel=Qwen3-8B2026.04 | 38 | — | |
| Base modelModel=Qwen3-4B-Instruct, Rollouts (N)=–2026.05 | 37.4 | — | |
| GPG-1.5BZero-shot=true, Model Category=Qwen 1.5B Models, Training Samples=7,0002025.05 | 36.7 | — | |
| PMPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 36.7 | — | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=-2, p-Scheduling Strategy=static2026.05 | 36.7 | — | |
| Llama3.3-70BReasoning Strategy=Direct Reasoning, Model Scale=> 32B Parameters2025.10 | 36.7 | — | |
| Auto-TIRReasoning Strategy=Tool Integrated, Backbone=Qwen2.5-7B2025.10 | 33.33 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BZero-shot=true, Model Category=Qwen 1.5B Models, Training Samples=–2025.05 | 33.3 | — | |
| AAPO-1.5BZero-shot=true, Model Category=Qwen 1.5B Models, Training Samples=7,0002025.05 | 33.3 | — | |
| GPG-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 33.3 | — | |
| GRPO-VPSBackbone=Qwen2.5-Math-7B, Supervision Type=With Process Supervision2026.04 | 31.7 | 3,829 | |
| GSPOBackbone=Qwen2.5-Math-7B, Supervision Type=Outcome Supervision Only2026.04 | 30.8 | 4,191 | |
| SimpleRL-Zero-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 30 | — | |
| Oat-Zero-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 30 | — | |
| AAPO-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 30 | — | |
| GRPOBackbone=Qwen2.5-Math-7B, Supervision Type=Outcome Supervision Only2026.04 | 30 | 4,730 | |
| GRPO w/ Skywork-7BBackbone=Qwen2.5-Math-7B, Supervision Type=With Process Supervision2026.04 | 30 | 4,697 | |
| HölderPO-1.5BModel Scale=1.5B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 30 | — | |
| DrGRPOBackbone=Qwen2.5-Math-7B, Supervision Type=Outcome Supervision Only2026.04 | 28.3 | 4,449 | |
| GRPO-1.5BZero-shot=true, Model Category=Qwen 1.5B Models, Training Samples=7,0002025.05 | 26.7 | — | |
| SimpleRL-Zero-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 26.7 | — | |
| SFTAccess=text2026.05 | 26.69 | — | |
| S-GRPOBackbone=Qwen2.5-Math-7B, Supervision Type=With Process Supervision2026.04 | 25.8 | 3,723 | |
| Qwen3-4B (student)Access=–2026.05 | 24.17 | — | |
| GPG-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 23.3 | — | |
| OpenReasoner-Zero-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=57K2025.05 | 20 | — | |
| DrGRPOBackbone=Qwen2.5-Math-1.5B, Supervision Type=Outcome Supervision Only2026.04 | 20 | 4,878 | |
| Oat-Zero-1.5BModel Scale=1.5B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 20 | — | |
| GMPO-1.5BModel Scale=1.5B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 20 | — | |
| Qwen2.5-Coder-32BReasoning Strategy=Direct Reasoning, Model Scale=> 32B Parameters2025.10 | 20 | — | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=322026.05 | 19.5 | — | |
| NUDGERLModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 19 | — | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=162026.05 | 18.8 | — | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 18.7 | — | |
| POPEModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 18.6 | — | |
| GRPOBackbone=Qwen2.5-Math-1.5B, Supervision Type=Outcome Supervision Only2026.04 | 18.3 | 4,655 | |
| Eurus-2-7B-PRIMEZero-shot=true, Model Category=Qwen 7B Models, Training Samples=230K + 150K2025.05 | 16.7 | — | |
| Qwen2.5-Math-1.5BModel Scale=1.5B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=Base2026.05 | 16.7 | — | |
| Qwen2.5-Math-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=Base2026.05 | 16.7 | — | |
| PRIME-Zero-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 16.7 | — | |
| Eurus-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 16.7 | — | |
| STEP-KTO M3Reasoning Strategy=Custom Training, Backbone=Llama-3.1-8B-Instruct2025.10 | 16.7 | — | |
| Qwen2.5-7BReasoning Strategy=Direct Reasoning, Model Scale=< 8B Parameters2025.10 | 16.67 | — | |
| SIGMAReasoning Strategy=Tool Integrated, Backbone=Qwen2.5-7B2025.10 | 16.67 | — | |
| GRPO-VPSBackbone=Qwen2.5-Math-1.5B, Supervision Type=With Process Supervision2026.04 | 15 | 4,278 | |
| Eurus-2-7B-PRIMEBackbone=Qwen2.5-Math-7B, Supervision Type=With Process Supervision2026.04 | 15 | 5,731 | |
| S-GRPOBackbone=Qwen2.5-Math-1.5B, Supervision Type=With Process Supervision2026.04 | 14.2 | 4,580 | |
| Base modelModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=–2026.05 | 13.4 | — | |
| Search-o1Reasoning Strategy=Tool Integrated, Backbone=Qwen2.5-7B2025.10 | 13.33 | — | |
| GSPOBackbone=Qwen2.5-Math-1.5B, Supervision Type=Outcome Supervision Only2026.04 | 13.3 | 5,146 | |
| OpenReasoner-Zero-7B @ 3kModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 13.3 | — | |
| OpenReasoner-Zero-7B @ 8kModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 13.3 | — | |
| GRPO w/ Skywork-1.5BBackbone=Qwen2.5-Math-1.5B, Supervision Type=With Process Supervision2026.04 | 11.7 | 4,738 | |
| BASEBackbone=Qwen2.5-Math-7B, Supervision Type=Outcome Supervision Only2026.04 | 10.8 | 5,390 | |
| OverRIDEModel=Qwen2.5-7B2026.04 | 10.8 | — | |
| VanillaModel=Qwen2.5-7B2026.04 | 10.7 | — | |
| Min-pModel=Qwen2.5-7B2026.04 | 10.3 | — | |
| Llama-3.2-1B-Instruct + AAPOZero-shot=true, Model Category=Llama 1B Models, Training Samples=8,5232025.05 | 10 | — | |
| Qwen2.5-Math-1.5B-InstructModel Scale=1.5B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=Instruct2026.05 | 10 | — | |
| ESampModel=Qwen2.5-7B2026.04 | 9.5 | — | |
| GPT-4oReasoning Strategy=Direct Reasoning, Model Scale=> 32B Parameters2025.10 | 9.3 | — | |
| FIREModel=Qwen2.5-7B2026.04 | 9.2 | — | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=642026.05 | 8.1 | — | |
| Llama-3.2-3B-Instruct + GPGZero-shot=true, Model Category=Llama 3B Models, Training Samples=8,5232025.05 | 6.7 | — |