Mathematical Reasoning on GSM8K (pass@1, maj@8, rm@8)
96.7pass@1Qwen2-Math-72B-Instruct
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen2-Math-72B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 96.7 | 97 | 96.7 | 65.7 | |
| Qwen2.5-Math-72B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 95.9 | 96 | 96.4 | 68.2 | |
| Qwen2.5-Math-72B-InstructReasoning Mode=Tool-Integrated Reasoning, Shots=0-shot2024.09 | 95.8 | 96.7 | 96.4 | 72.6 | |
| Qwen2.5-Math-7B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 95.2 | 96.7 | 97.9 | 62.9 | |
| Qwen2.5-Math-7B-InstructReasoning Mode=Tool-Integrated Reasoning, Shots=0-shot2024.09 | 94.6 | 96.4 | 97.6 | 67.4 | |
| OpenMath-Nemotron-14B + iGRPOBackbone=Nemotron-14B, Fine-tuning=iGRPO2026.02 | 94.16 | — | — | — | |
| Llama-3.1-70B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 94.1 | — | — | 56.9 | |
| OpenMath-Nemotron-14BBackbone=Nemotron-14B2026.02 | 94.01 | — | — | — | |
| w/o MergingModels=Math2026.02 | 93.73 | — | — | — | |
| GC2POBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 93.6 | — | — | — | |
| w/o MergingModels=Code2026.02 | 93.46 | — | — | — | |
| SCF-RKLModels=Fuse2026.02 | 93.42 | — | — | — | |
| Qwen2-72B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 93.2 | — | — | 59 | |
| SetPO+DAPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=DAPO2026.02 | 93 | — | — | — | |
| GPT-4o-2024-08-06Reasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 92.9 | — | — | 62 | |
| GVPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.9 | — | — | — | |
| L2T-GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.9 | — | — | — | |
| GCPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.6 | — | — | — | |
| GSPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 92.4 | — | — | — | |
| Dr.GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.3 | — | — | — | |
| DAPOBase Model=Qwen2.5-Math-7B2026.02 | 92.2 | — | — | — | |
| SetPO+GRPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=GRPO2026.02 | 92.2 | — | — | — | |
| MRTBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.2 | — | — | — | |
| Internlm2-math-plus-mixtral8x7BReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 92.1 | — | — | 51.8 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92.1 | — | — | — | |
| ReST-MCTSBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 92 | — | — | — | |
| M2POModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 91.8 | — | — | — | |
| SetPO+GSPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=GSPO2026.02 | 91.6 | — | — | — | |
| GC2POBackbone=DeepScaleR-1.5B-Preview2026.02 | 91.6 | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7BFine-tuning=None2026.02 | 91.6 | — | — | — | |
| length penaltyBackbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91.1 | — | — | — | |
| L2T-GRPOBackbone=DeepScaleR-1.5B-Preview2026.02 | 90.9 | — | — | — | |
| NuminaMath-72B-CoTReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 90.8 | — | — | 54 | |
| R1-zero-DivBase Model=Qwen2.5-Math-7B2026.02 | 90.6 | — | — | — | |
| GCPOBackbone=DeepScaleR-1.5B-Preview2026.02 | 90.5 | — | — | — | |
| GVPOBackbone=DeepScaleR-1.5B-Preview2026.02 | 90.4 | — | — | — | |
| MinPROModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 90.2 | — | — | — | |
| MinPROModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 90.1 | — | — | — | |
| Dr.GRPOBackbone=DeepScaleR-1.5B-Preview2026.02 | 90 | — | — | — | |
| Qwen2-Math-7B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 89.9 | 93.1 | 95.1 | 58.5 | |
| ReST-MCTSBackbone=DeepScaleR-1.5B-Preview2026.02 | 89.9 | — | — | — | |
| MRTBackbone=DeepScaleR-1.5B-Preview2026.02 | 89.9 | — | — | — | |
| DeepScaleR-1.5B-PreviewFine-tuning=None2026.02 | 89.6 | — | — | — | |
| GSPOBase Model=Qwen2.5-Math-7B2026.02 | 89.5 | — | — | — | |
| GRPOBackbone=DeepScaleR-1.5B-Preview2026.02 | 89.4 | — | — | — | |
| GRPOBase Model=Qwen2.5-Math-7B2026.02 | 89.1 | — | — | — | |
| M2POModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 88.7 | — | — | — | |
| length penaltyBackbone=DeepScaleR-1.5B-Preview2026.02 | 88.5 | — | — | — | |
| DeepSeekMath-7B-RLReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 88.2 | — | — | 46.6 | |
| Internlm2-math-plus-20BReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 87.9 | — | — | 48.7 | |
| DeepSeek-Coder-V2-Lite-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 87.6 | — | — | 52.7 | |
| GSPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 87.3 | — | — | — | |
| CISPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 87 | — | — | — | |
| CISPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 86.7 | — | — | — | |
| GC2POBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 86.5 | — | — | — | |
| GRPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 86.2 | — | — | — | |
| Qwen2-7B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 85.7 | — | — | 44.1 | |
| L2T-GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 85.5 | — | — | — | |
| GCPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 85.3 | — | — | — | |
| GVPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 85 | — | — | — | |
| Dr.GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 85 | — | — | — | |
| Mathstral-7B-v0.1Reasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 84.9 | — | — | 46.1 | |
| Qwen2.5-Math-1.5B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 84.8 | 88.6 | 92.5 | 56.9 | |
| ReST-MCTSBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.8 | — | — | — | |
| MRTBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.7 | — | — | — | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.5 | — | — | — | |
| Qwen2-Math-1.5B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 84.2 | 88.6 | 92.7 | 53.3 | |
| Internlm2-math-plus-7BReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 84 | — | — | 45.1 | |
| Qwen2.5-Math-1.5B-InstructReasoning Mode=Tool-Integrated Reasoning, Shots=0-shot2024.09 | 83.7 | 90 | 93.3 | 56.9 | |
| DeepSeek-R1-Distill-Qwen-1.5BFine-tuning=None2026.02 | 83.4 | — | — | — | |
| length penaltyBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 82.6 | — | — | — | |
| GRPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 80.4 | — | — | — | |
| Llama-3.1-8B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 76.6 | — | — | 41.9 | |
| NuminaMath-7B-CoTReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 75.4 | — | — | 45 | |
| w/o MergingModel=Meta-Llama-3-8B-Instruct (base0)2026.02 | 73.35 | — | — | — | |
| SCEModels=Fuse2026.02 | 72.65 | — | — | — | |
| Dare Ties MergingModels=Fuse2026.02 | 71.74 | — | — | — | |
| SCF-RKLModel=Fused2026.02 | 71.68 | — | — | — | |
| Task ArithmeticModels=Fuse2026.02 | 71.49 | — | — | — | |
| Dare Task ArithmeticModels=Fuse2026.02 | 71.44 | — | — | — | |
| Ties MergingModels=Fuse2026.02 | 70.74 | — | — | — | |
| BaseModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 70.3 | — | — | — | |
| BaseModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 68.4 | — | — | — | |
| Qwen2-1.5B-InstructReasoning Mode=Chain-of-Thought, Shots=0-shot2024.09 | 64.1 | — | — | 25 | |
| Dare Ties MergingModels=Fuse2026.02 | 58.53 | — | — | — | |
| SCF-RKLModels=Fuse2026.02 | 58.38 | — | — | — | |
| w/o MergingModel=MAmmoTH2-8B-Plus (base1)2026.02 | 56.54 | — | — | — | |
| Qwen2.5-Math-7BModel Type=Base Model2026.02 | 53.4 | — | — | — | |
| Task ArithmeticModels=Fuse2026.02 | 52.08 | — | — | — | |
| Dare Task ArithmeticModels=Fuse2026.02 | 52.08 | — | — | — | |
| Dare Ties MergingModel=Fused2026.02 | 50.11 | — | — | — | |
| Task ArithmeticModel=Fused2026.02 | 49.56 | — | — | — | |
| Mistral-7B-Instruct-v0.2Models=base02026.02 | 46.76 | — | — | — | |
| LESASFT=true, training_steps=6k2025.02 | 37.14 | — | — | 31.57 | |
| SOLARSFT=true, training_steps=6k2025.02 | 33.45 | — | — | 26.47 | |
| LLaMA ProSFT=true, training_steps=6k2025.02 | 21.95 | — | — | 24.38 | |
| MathCoder2-Mistral-7BModels=base12026.02 | 14.95 | — | — | — | |
| Ties MergingModels=Fuse2026.02 | 14.75 | — | — | — | |
| SCEModels=Fuse2026.02 | 14.39 | — | — | — | |
| Ties MergingModel=Fused2026.02 | 6.31 | — | — | — |