Mathematical Reasoning on MATH (Accuracy)
94.2AccuracySelf-Reminder
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Self-ReminderBackbone=Qwen3-14B2026.01 | 94.2 | — | |
| InstructBackbone=Qwen3-14B2026.01 | 92.8 | — | |
| Self-GuardBackbone=Qwen3-14B2026.01 | 92.8 | — | |
| ELPO-4BStarting Checkpoint=Inst, Backbone Architecture Family=Qwen3-4B2026.02 | 92.8 | — | |
| ReasoningGuardBackbone=Qwen3-8B2026.01 | 92.4 | — | |
| SafeKeyBackbone=Qwen3-14B2026.01 | 92.4 | — | |
| ReasoningGuardBackbone=Qwen3-14B2026.01 | 92.4 | — | |
| InstructBackbone=Qwen3-8B2026.01 | 92.2 | — | |
| STAR-1Backbone=Qwen3-14B2026.01 | 92 | — | |
| DemyAgent-4BStarting Checkpoint=Inst, Backbone Architecture Family=Qwen3-4B2026.02 | 91.6 | — | |
| ELPO-7BStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 91.2 | — | |
| Self-ReminderBackbone=Qwen3-8B2026.01 | 91 | — | |
| DemyAgentStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 90.8 | — | |
| Self-GuardBackbone=Qwen3-4B2026.01 | 90.6 | — | |
| SafeKeyBackbone=Qwen3-8B2026.01 | 90.6 | — | |
| STAR-1Backbone=Qwen3-8B2026.01 | 90.4 | — | |
| SafeChainBackbone=Qwen3-14B2026.01 | 90.4 | — | |
| CIRStarting Checkpoint=Math, Backbone Architecture Family=Qwen2.5-7B2026.02 | 90.4 | — | |
| ReasoningGuardBackbone=Qwen3-4B2026.01 | 90.2 | — | |
| Self-GuardBackbone=Qwen3-8B2026.01 | 90.2 | — | |
| AEPOStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 90 | — | |
| Self-ReminderBackbone=Qwen3-4B2026.01 | 89.8 | — | |
| InstructBackbone=Qwen3-4B2026.01 | 89.6 | — | |
| STAR-1Backbone=Qwen3-4B2026.01 | 88.9 | — | |
| GRPO w/Clip-higher + LIEBackbone=Qwen3-4B-Base, Variant=Clip-higher, Strategy=Length-Incentivized Exploration2026.02 | 88.8 | — | |
| Reinforce++Starting Checkpoint=Instruct, Backbone Architecture Family=Qwen2.5-7B2026.02 | 88.8 | — | |
| DAPOStarting Checkpoint=Instruct, Backbone Architecture Family=Qwen2.5-7B2026.02 | 88.8 | — | |
| ARPOStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 88.8 | — | |
| PADsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 88.7 | — | |
| PADsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 88.41 | — | |
| PADsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 88.4 | — | |
| GSPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 88.4 | — | |
| GRPOStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 87.8 | — | |
| ToRLStarting Checkpoint=Math-Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 87.8 | — | |
| RSFTsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 87.64 | — | |
| GIGPOStarting Checkpoint=Inst, Backbone Architecture Family=Qwen2.5-7B2026.02 | 87.6 | — | |
| RSFTsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 87.19 | — | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-8, Bound=Upper Bound2026.01 | 87 | — | |
| DAsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 86.41 | — | |
| GRPO w/Clip-higherBackbone=Qwen3-4B-Base, Variant=Clip-higher2026.02 | 86.4 | — | |
| DAsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 86.2 | — | |
| SafeChainBackbone=Qwen3-8B2026.01 | 86.2 | — | |
| DAsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 86.05 | — | |
| RSFTsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 85.97 | — | |
| SafeChainBackbone=Qwen3-4B2026.01 | 85.6 | — | |
| GSPOBackbone=Qwen3-4B-Base2026.02 | 85.2 | — | |
| GRPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 85 | — | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-7B-Instruct2026.01 | 84.2 | — | |
| Qwen2.5-Math-7B-NAITTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 82.1 | — | |
| Majority Vote@8Policy Model=Qwen2.5-14B-Instruct, Strategy=Majority Voting, Samples=82026.01 | 81.8 | — | |
| Qwen2.5-Math-PRM-7BTraining Samples=1500K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 81.6 | — | |
| Qwen3-4BStarting Checkpoint=Inst, Reasoning Strategy=TIR Reasoning, Backbone Architecture Family=Qwen3-4B2026.02 | 81.5 | — | |
| FlowSteerFramework=Ours, Backbone=4o-mini2026.02 | 81.25 | — | |
| RMoAModel=GPT-4o2025.05 | 81.16 | — | |
| EurusPRM-Stage2Training Samples=693K, Aggregation Method=Sum, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 80.6 | — | |
| Skywork-PRM-Qwen2.5-7BTraining Samples=N/A, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 80.6 | — | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=-, Aggregation Method=Mean2026.01 | 80.6 | — | |
| GRPOBackbone=Qwen3-4B-Base2026.02 | 80.4 | — | |
| Qwen3-4BStarting Checkpoint=Inst, Reasoning Strategy=Self-Contained Reasoning, Backbone Architecture Family=Qwen3-4B2026.02 | 80.4 | — | |
| Qwen2.5-Math-7B-MCRDTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 80.2 | — | |
| Qwen2.5-Math-7B-NAITPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 80.1 | — | |
| MoAModel=GPT-4o2025.05 | 80.08 | — | |
| RLHFlow-PRM-Mistral-8BTraining Samples=273K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 80 | — | |
| Qwen2.5-Math-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=1500K, Aggregation Method=Mean2026.01 | 79.8 | — | |
| GSPO (No Constraint)Backbone=Qwen2.5-7B2026.01 | 79.8 | — | |
| EurusPRM-Stage1Policy Model=Qwen2.5-7B-Instruct, Training Samples=463K, Aggregation Method=Min-Max2026.01 | 79.6 | — | |
| RLHFlow-PRM-DeepSeek-8BTraining Samples=253K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 79.4 | — | |
| EurusPRM-Stage1Training Samples=463K, Aggregation Method=Min, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 79.4 | — | |
| EurusPRM-Stage2Policy Model=Qwen2.5-7B-Instruct, Training Samples=693K, Aggregation Method=Sum2026.01 | 79.2 | — | |
| Majority Vote@8Policy Model=Qwen2.5-7B-Instruct2026.01 | 78.6 | — | |
| Qwen2.5-7BStarting Checkpoint=Inst, Reasoning Strategy=TIR Reasoning, Backbone Architecture Family=Qwen2.5-7B2026.02 | 78.2 | — | |
| SMoAModel=GPT-4o2025.05 | 78.08 | — | |
| GreedyPolicy Model=Qwen2.5-14B-Instruct, Strategy=Greedy Search2026.01 | 78 | — | |
| CARE-GSPOBackbone=Qwen2.5-7B2026.01 | 77.6 | — | |
| Qwen2.5-Math-7B-MCRDPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 77.5 | — | |
| RMoAModel=Qwen2.5-7B-Instruct2025.05 | 77.2 | — | |
| SMoAModel=Qwen2.5-7B-Instruct2025.05 | 76.98 | — | |
| RLHFlow-PRM-DeepSeek-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=253K, Aggregation Method=Mean2026.01 | 76.8 | — | |
| GPT-4oModel=GPT-4o2025.05 | 76.6 | — | |
| SafeKeyBackbone=Qwen3-4B2026.01 | 76.6 | — | |
| Math-Shepherd-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=445K, Aggregation Method=Mean2026.01 | 76.6 | — | |
| RLHFlow-PRM-Mistral-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=273K, Aggregation Method=Mean2026.01 | 76.6 | — | |
| Router-R1Framework=Agent+RL, Backbone=4o-mini2026.02 | 76.56 | — | |
| Math-Shepherd-PRM-7BTraining Samples=445K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 76.2 | — | |
| Qwen2.5-Math-7B-MCPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 76.2 | — | |
| DAPO (No Constraint)Backbone=Qwen2.5-7B2026.01 | 76.1 | — | |
| Qwen2.5-Math-7B-MCTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 75.9 | — | |
| Qwen2.5-7BStarting Checkpoint=Inst, Reasoning Strategy=Self-Contained Reasoning, Backbone Architecture Family=Qwen2.5-7B2026.02 | 75.5 | — | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Pass@82026.01 | 75.4 | — | |
| MoAModel=Qwen2.5-7B-Instruct2025.05 | 75.28 | — | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct2025.05 | 74.94 | — | |
| GreedyPolicy Model=Qwen2.5-7B-Instruct2026.01 | 74 | — | |
| Majority Vote@8Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Majority Vote@82026.01 | 74 | — | |
| CARE-DAPOBackbone=Qwen2.5-7B2026.01 | 74 | — | |
| Alpha-steerBackbone=Qwen3-4B2026.01 | 73.4 | — | |
| Alpha-steerBackbone=Qwen3-14B2026.01 | 73.4 | — | |
| GRPO (No Constraint)Backbone=Qwen2.5-7B2026.01 | 72.4 | — | |
| OrchestratorFramework=Agent+RL, Backbone=4o-mini2026.02 | 72.26 | — | |
| AgentflowFramework=Agent+RL, Backbone=4o-mini2026.02 | 71.87 | — | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Aggregation Method=Mean2026.01 | 70.6 | — |