Mathematical Reasoning on MATH 500 (Top-1 Accuracy)
96.4Top-1 AccuracyPADD
Evaluation Results
| Method | Links | |
|---|---|---|
| PADDModel Family=Qwen Family2026.06 | 96.4 | |
| SFT+RLBackbone=Qwen3-32B, Tokens=18322026.04 | 95 | |
| Abstract-CoT (Warm-up + RL)Backbone=Qwen3-32B, Tokens=1672026.04 | 94.6 | |
| SMCSMaximum output tokens=32,7682025.07 | 94.5 | |
| RSPOModel Family=Qwen Family2026.06 | 94.1 | |
| GSPOModel Family=Qwen Family2026.06 | 93.7 | |
| SuCoModel Scale=7B2026.06 | 93.6 | |
| SFT (CoT)Backbone=Qwen3-32B, Tokens=17062026.04 | 93.4 | |
| PMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 93.4 | |
| LHRMsModel Scale=7B2026.06 | 93 | |
| Qwen3-32BMaximum output tokens=32,7682025.07 | 92.8 | |
| AdaptThinkModel Scale=7B2026.06 | 92.8 | |
| HölderPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 92.6 | |
| GLM-Z1-32B-0414Maximum output tokens=32,7682025.07 | 92.2 | |
| EXAONE-Deep-32BMaximum output tokens=32,7682025.07 | 92.2 | |
| S-GRPOModel Scale=7B2026.06 | 92.2 | |
| Online KDModel Family=Qwen Family2026.06 | 92.1 | |
| QwQ-32BMaximum output tokens=32,7682025.07 | 91.8 | |
| AdaCoTModel Scale=7B2026.06 | 91.8 | |
| GMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=p → 02026.05 | 91.4 | |
| Vanilla-GRPOModel Family=Qwen Family2026.06 | 91.4 | |
| Teacher (GRPO)Model Family=Qwen Family2026.06 | 91.3 | |
| Stepwise InternalizationBackbone=Qwen3-32B, Tokens=1632026.04 | 90.6 | |
| BaseModel Family=Qwen Family2026.06 | 90.5 | |
| Abstract-CoT (Warm-up)Backbone=Qwen3-32B, Tokens=1952026.04 | 90.2 | |
| Hi-CoT (format-relaxed)Model=Qwen3-4B-Instruct-25072026.03 | 90 | |
| Dr.GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 89.6 | |
| SFT (no CoT)Backbone=Qwen3-32B, Tokens=4272026.04 | 89 | |
| GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=12026.05 | 89 | |
| DeepSeek-R1-DistillModel Scale=7B2026.06 | 89 | |
| CoTModel=Qwen3-4B-Instruct-25072026.03 | 88.8 | |
| Hi-CoTModel=Qwen3-4B-Instruct-25072026.03 | 88.8 | |
| DenoiseRL-DAPOBase Model=Qwen3-8B-Base2026.05 | 88.2 | |
| MTLModel Scale=14B2026.05 | 87.8 | |
| GRPOBase Model=Qwen3-8B-Base2026.05 | 87.8 | |
| Plan-and-SolveModel=Qwen3-4B-Instruct-25072026.03 | 87.4 | |
| GMMSize=70B, Selection=Best-of-1002026.04 | 87.26 | |
| EGRSDBackbone=Qwen3-8B2026.05 | 87.2 | |
| DenoiseRL-GRPOBase Model=Qwen3-8B-Base2026.05 | 87.2 | |
| OPSDBackbone=Qwen3-8B2026.05 | 87.11 | |
| DAPOBase Model=Qwen3-8B-Base2026.05 | 87 | |
| CL-EGRSDBackbone=Qwen3-8B2026.05 | 86.95 | |
| DCLSize=70B, Selection=Best-of-1002026.04 | 86.93 | |
| CRISPBackbone=Qwen3-8B2026.05 | 86.93 | |
| Baseline (no-train)Backbone=Qwen3-8B2026.05 | 86.92 | |
| BaselineBackbone=Qwen3-32B, Tokens=12782026.04 | 86.8 | |
| SuCoModel Scale=1.5B2026.06 | 86.8 | |
| DDRLBase Model=Qwen2.5-Math-7B2026.04 | 86.7 | |
| SFTBackbone=Qwen3-8B2026.05 | 86.6 | |
| GRPOBackbone=Qwen3-8B2026.05 | 86.6 | |
| StandardModel=Qwen3-4B-Instruct-25072026.03 | 86.2 | |
| DeepSeek-R1-Distill-Llama-70BMaximum output tokens=32,7682025.07 | 86 | |
| GPT-4.1(2025-04-14)Maximum output tokens=32,7682025.07 | 85.8 | |
| DenoiseRL-GRPOBase Model=Qwen3-4B-Base2026.05 | 85.8 | |
| DeepSeek-R1-Distill-Qwen-32BMaximum output tokens=32,7682025.07 | 85.6 | |
| Qwen2.5-Math-7B-Instructhint_setting=w/ LLM Hints2026.04 | 85.14 | |
| DeepSeek-R1-Distill-Qwen-7Bhint_setting=w/ SLM Hints2026.04 | 85.14 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=-1, p-Scheduling Strategy=static2026.05 | 85 | |
| GRPOBackbone=Qwen2.5-Math-7B2026.04 | 84.9 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=-2, p-Scheduling Strategy=static2026.05 | 84.6 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=p → 0, p-Scheduling Strategy=dynamic2026.05 | 84.6 | |
| DenoiseRL-DAPOBase Model=Qwen3-4B-Base2026.05 | 84.6 | |
| SHEARBackbone=Qwen2.5-14B-Base2026.04 | 84.5 | |
| Abstract-CoT (RL-only)Backbone=Qwen3-32B, Tokens=1372026.04 | 84.4 | |
| GPT-o3-mini(2025-01-31)Maximum output tokens=32,7682025.07 | 84.4 | |
| LHRMsModel Scale=1.5B2026.06 | 84.4 | |
| S-GRPOModel Scale=1.5B2026.06 | 84.2 | |
| Gemma-3-27b-itMaximum output tokens=32,7682025.07 | 84 | |
| POLISModel=Gemma3-4b2025.07 | 83.9 | |
| PRIME-Zero-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 83.8 | |
| Eurus-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 83.8 | |
| PMPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 83.8 | |
| DAPOBase Model=Qwen3-4B-Base2026.05 | 83.8 | |
| AdaptThinkModel Scale=1.5B2026.06 | 83.8 | |
| GRPOBase Model=Qwen3-4B-Base2026.05 | 83.6 | |
| DALAAgent Category=Multi-Agent2025.11 | 83.42 | |
| TTRLBase Model=Qwen2.5-Math-7B2026.04 | 83.4 | |
| Entropy adv.Backbone=Qwen2.5-Math-7B2026.04 | 83.4 | |
| GRPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=12026.05 | 83.4 | |
| SHEARBackbone=Qwen2.5-Math-7B2026.04 | 83.3 | |
| GRPOBackbone=Qwen2.5-14B-Base2026.04 | 83.3 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=1, p-Scheduling Strategy=static2026.05 | 83.2 | |
| AdaCoTModel Scale=1.5B2026.06 | 83.2 | |
| PRM(Reshape adv.)Backbone=Qwen2.5-Math-7B2026.04 | 83.1 | |
| Dense-GRPOModel Family=Qwen Family2026.06 | 83.1 | |
| DisCOBackbone=Qwen3-4B-Base2026.05 | 83 | |
| AgentPrune-RAgent Category=Multi-Agent2025.11 | 82.81 | |
| Hi-CoTModel=DeepSeek-R1-Distill-Qwen-32B2026.03 | 82.6 | |
| Pause TokensBackbone=Qwen3-32B, Tokens=1562026.04 | 82.6 | |
| Dr. GRPOBackbone=Qwen3-4B-Base2026.05 | 82.6 | |
| PRM(PURE)Backbone=Qwen2.5-Math-7B2026.04 | 82.5 | |
| DeepSeek-R1-Distill-Qwen-7Bhint_setting=w/ LLM Hints2026.04 | 82.43 | |
| OpenReasoner-Zero-7B @ 8kModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 82.4 | |
| Entropy adv.Backbone=Qwen2.5-14B-Base2026.04 | 82.2 | |
| OPOBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 82.2 | |
| Re-Schedule_sigmoidBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@32, Weighting Scheme=sigmoid2025.10 | 82.2 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 82.2 | |
| ConSPOBackbone=Qwen3-4B-Base2026.05 | 82.2 | |
| GMPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=p → 02026.05 | 82 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=2, p-Scheduling Strategy=static2026.05 | 82 |