Mathematical Reasoning on AMC
97.5AccuracyGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Evaluation Protocol=Closed-Source2026.01 | 97.5 | |
| Qwen3-235B-A22BModel Category=Large Language Models, Model Scale=235B-A22B, Reasoning Strategy=Thinking2025.12 | 93.98 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Thinking2025.12 | 93.98 | |
| FIGR2025.12 | 93.98 | |
| GLM-4.5VModel Category=Large Vision-Language Models, Model Scale=108B2025.12 | 87.95 | |
| Text-only RLMode=Text-only, Reasoning Strategy=RL2025.12 | 87.95 | |
| Qwen3-VL-32B-InstructModel Category=Large Vision-Language Models, Model Scale=32B, Reasoning Strategy=Instruct2025.12 | 84.34 | |
| TTRL-GTModel=Qwen2.5-7B-Base2026.03 | 82.61 | |
| GPT-4.1Evaluation Protocol=Closed-Source2026.01 | 82.5 | |
| ATLAS (cluster)Evaluation Protocol=In-Distribution2026.01 | 82.5 | |
| Qwen3-VL-8B-InstructModel Category=Large Vision-Language Models, Model Scale=8B, Reasoning Strategy=Instruct2025.12 | 81.93 | |
| PMPO-7B (Ours) [R1-Distill]Backbone=DeepSeek-R1-Distill-Qwen-7B, Optimization Algorithm=PMPO2026.01 | 79.5 | |
| GMPO-7B [R1-Distill]Backbone=DeepSeek-R1-Distill-Qwen-7B, Optimization Algorithm=GMPO2026.01 | 78.3 | |
| DCPOBackbone=Qwen2.5-math-7B, Optimization Method=DCPO, Sampling Strategy=x322026.01 | 76.3 | |
| Dr.GRPO-7B [R1-Distill]Backbone=DeepSeek-R1-Distill-Qwen-7B, Optimization Algorithm=Dr.GRPO2026.01 | 74.7 | |
| BAPORollouts=733k, Type=off, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 72.74 | |
| R-Few (5%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=5%2025.12 | 72.3 | |
| DistriTTRL-GMMModel=Qwen3-8B2026.03 | 70.31 | |
| DAPORollouts=1921k, Type=on, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 70.08 | |
| SPICEBackbone=Qwen3-8B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 70 | |
| DCPOBackbone=Qwen2.5-7B, Optimization Method=DCPO, Sampling Strategy=x322026.01 | 69.9 | |
| R-Few (1%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 69.3 | |
| TTRL-WSCModel=Qwen3-8B2026.03 | 68.8 | |
| PMPO-7B (Ours)Backbone=Qwen2.5-Math-7B, Optimization Algorithm=PMPO2026.01 | 68.7 | |
| TTRL-GTModel=Qwen3-8B2026.03 | 67.87 | |
| ATLAS (RL)Evaluation Protocol=Out-of-Distribution2026.01 | 67.5 | |
| CAPO (RLOO)Backbone=Qwen2.5-7B-Math2025.12 | 67.5 | |
| GRPORollouts=677k, Type=on, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 67.47 | |
| DistriTTRL-WSCModel=Qwen3-8B2026.03 | 66.57 | |
| GSPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 66.2 | |
| TTRLModel=Qwen3-8B2026.03 | 65.95 | |
| GRPO w/Clip-higher + LIEBackbone=Qwen3-4B-Base, Variant=Clip-higher, Strategy=Length-Incentivized Exploration2026.02 | 65.5 | |
| MoPPSRollouts=737k, Type=on, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 65.29 | |
| DEEP-GRPO-7BParameter size=7B2026.02 | 65.1 | |
| GRPO (v = 5)Rollouts=677k, Type=off, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 65.09 | |
| Remix-GRPORollouts=N/A, Type=off, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 65.06 | |
| GPG-7BBackbone=Qwen2.5-Math-7B, Optimization Algorithm=GPG2026.01 | 65 | |
| GPG-7BParameter size=7B2026.02 | 65 | |
| CAPO (GRPO)Backbone=Qwen2.5-7B-Math2025.12 | 65 | |
| General-ReasonerBackbone=Qwen3-8B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 64.8 | |
| RePORollouts=677k, Type=off, Backbone=DeepSeek R1 Distill Qwen 1.5B2026.02 | 64.76 | |
| DeepSeek R1 Distill Qwen 1.5BRollouts=N/A, Type=N/A2026.02 | 62.9 | |
| R-ZeroBackbone=Qwen3-8B-Base, Methodology=R-Zero, Supervision Ratio=None2025.12 | 62.8 | |
| GSPOBackbone=Qwen3-4B-Base2026.02 | 62.7 | |
| PRIME-Zero-7BBackbone=Qwen2.5-Math-7B, Optimization Algorithm=PRIME-Zero2026.01 | 62.7 | |
| Eurus-7BBackbone=Mistral-7B, Optimization Algorithm=Eurus2026.01 | 62.7 | |
| Dr.GRPO-7BBackbone=Qwen2.5-Math-7B, Optimization Algorithm=Dr.GRPO2026.01 | 62.7 | |
| PRIME-Zero-7BParameter size=7B2026.02 | 62.7 | |
| Eurus-7BParameter size=7B2026.02 | 62.7 | |
| Oat-Zero-7B (Dr. GRPO)Parameter size=7B2026.02 | 62.7 | |
| Absolute ZeroBackbone=Qwen3-8B-Base, Methodology=Absolute Zero, Supervision Ratio=None2025.12 | 62.5 | |
| Gemini2.5-ProEvaluation Protocol=Closed-Source2026.01 | 62.5 | |
| RouterDCEvaluation Protocol=In-Distribution2026.01 | 62.5 | |
| CAPO (GRPO)Backbone=Qwen2.5-1.5B-Math2025.12 | 62.5 | |
| CHORD-ϕmechanism=dual-control2025.08 | 62.5 | |
| GRPO w/Clip-higherBackbone=Qwen3-4B-Base, Variant=Clip-higher2026.02 | 61.8 | |
| Base ModelBackbone=Qwen3-8B-Base, Methodology=Base, Supervision Ratio=None2025.12 | 61.5 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Non-Thinking2025.12 | 61.45 | |
| GMPO-7BBackbone=Qwen2.5-Math-7B, Optimization Algorithm=GMPO2026.01 | 61.4 | |
| DistriTTRL-GMMModel=Qwen2.5-Math-7B2026.03 | 61.37 | |
| CHORD-µtransition=smooth2025.08 | 60.8 | |
| GRPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 60.5 | |
| SimpleRL-Zero-7BBackbone=Qwen2.5-Math-7B, Optimization Algorithm=SimpleRL-Zero2026.01 | 60.2 | |
| SimpleRL-Zero-7BParameter size=7B2026.02 | 60.2 | |
| General-ReasonerBackbone=Qwen3-4B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 60 | |
| TTRL-GTModel=Qwen2.5-Math-7B2026.03 | 58.43 | |
| SFT-best + RLSFT intensity=thorough, RL=true2025.08 | 58.4 | |
| TTRL-WSCModel=Qwen2.5-Math-7B2026.03 | 58.28 | |
| DistriTTRL-WSCModel=Qwen2.5-Math-7B2026.03 | 57.68 | |
| SPICEBackbone=Qwen3-4B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 57.5 | |
| CAPO (PPO)Backbone=Qwen2.5-7B-Math2025.12 | 57.5 | |
| CAPO (PPO)Backbone=Qwen2.5-1.5B-Math2025.12 | 57.5 | |
| CAPO (RLOO)Backbone=Qwen2.5-1.5B-Math2025.12 | 57.5 | |
| DistriTTRL-WSCModel=Qwen2.5-7B-Base2026.03 | 57.23 | |
| ∇-ReasonerBackbone=Qwen-2.5-7B-Instruct, Nmax=82026.03 | 56.8 | |
| DistriTTRL-GMMModel=Qwen2.5-7B-Base2026.03 | 56.48 | |
| TTRLModel=Qwen2.5-Math-7B2026.03 | 56.18 | |
| BoNBackbone=Qwen-2.5-7B-Instruct, N=82026.03 | 55.9 | |
| TPOBackbone=Qwen-2.5-7B-Instruct2026.03 | 55.9 | |
| SFT-bestSFT intensity=thorough2025.08 | 55.9 | |
| DAREBackbone=Qwen3-1.7B, Method Category=Test-time scaling methods2026.01 | 55.7 | |
| SCBackbone=Qwen-2.5-7B-Instruct, N=82026.03 | 55.5 | |
| GRPOBackbone=Qwen3-4B-Base2026.02 | 55.2 | |
| RLOOBackbone=Qwen2.5-7B-Math2025.12 | 55 | |
| CAPO (Reinforce++)Backbone=Qwen2.5-7B-Math2025.12 | 55 | |
| RAPBackbone=Qwen-2.5-7B-Instruct2026.03 | 54.6 | |
| OpenReasoner-Zero-7B@8kBackbone=Qwen2.5-Math-7B, Optimization Algorithm=OpenReasoner-Zero, Sampling Budget=8k2026.01 | 54.2 | |
| OpenReasoner-Zero-7B @ 8kParameter size=7B, Inference budget=8k2026.02 | 54.2 | |
| TTRL-WSCModel=Qwen2.5-7B-Base2026.03 | 54.07 | |
| SASR2025.08 | 54 | |
| GRPOBackbone=Qwen3-1.7B, Method Category=RL methods on training dataset2026.01 | 53.2 | |
| Dr.GRPO-1.5BBackbone=Qwen2.5-Math-1.5B, Optimization Algorithm=Dr.GRPO2026.01 | 53 | |
| GMPO-1.5BBackbone=Qwen2.5-Math-1.5B, Optimization Algorithm=GMPO2026.01 | 53 | |
| RLPRBackbone=Qwen3-1.7B, Method Category=Test-time scaling methods2026.01 | 53 | |
| Oat-Zero-1.5B (Dr. GRPO)Parameter size=1.5B2026.02 | 53 | |
| TTRLBackbone=Qwen3-1.7B, Method Category=Test-time scaling methods2026.01 | 52.9 | |
| REINFORCE++Backbone=Qwen3-1.7B, Method Category=RL methods on training dataset2026.01 | 52.8 | |
| GRPOBackbone=Qwen-2.5-7B2026.03 | 52.8 | |
| LUFFY2025.08 | 52.8 | |
| R-Few (1%)Backbone=Qwen3-4B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 52.7 |