Mathematical Reasoning on MATH (Pass@1)
92.71Pass@1Thinker-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Thinker-7BModel Scale=7B2026.01 | 92.71 | |
| SATURN-7BModel Scale=7B2026.01 | 92.65 | |
| R³-7BModel Scale=7B2026.01 | 92.55 | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B2026.01 | 91.45 | |
| R³-1.5BModel Scale=1.5B2026.01 | 89.27 | |
| FASTCURL-1.5B-V2Model Scale=1.5B2026.01 | 89.25 | |
| Phi-4-mini + Mistral3-3BTraining=CORE2026.01 | 88.42 | |
| DeepScaleR-1.5B-PreviewModel Scale=1.5B2026.01 | 87.36 | |
| LENS_GRPOBase Model=Qwen3-8B-Base, RL Method=LENS, Rollout Configuration=Standard2026.01 | 86 | |
| GRPO_extendedBase Model=Qwen3-8B-Base, RL Method=GRPO, Rollout Configuration=2x Rollouts2026.01 | 85.2 | |
| DAPOBase Model=Qwen3-8B-Base, RL Method=DAPO, Rollout Configuration=Standard2026.01 | 85 | |
| DAPO_extendedBase Model=Qwen3-8B-Base, RL Method=DAPO, Rollout Configuration=2x Rollouts2026.01 | 85 | |
| GRESO_extendedBase Model=Qwen3-8B-Base, RL Method=GRESO, Rollout Configuration=2x Rollouts2026.01 | 84.6 | |
| GRESOBase Model=Qwen3-8B-Base, RL Method=GRESO, Rollout Configuration=Standard2026.01 | 84.2 | |
| GRPOBase Model=Qwen3-8B-Base, RL Method=GRPO, Rollout Configuration=Standard2026.01 | 84 | |
| STILL-3-1.5BModel Scale=1.5B2026.01 | 83.89 | |
| LENS_GRPOBase Model=Qwen3-4B-Base, RL Method=LENS, Rollout Configuration=Standard2026.01 | 83.2 | |
| SimpleRL-7BModel Scale=7B2026.01 | 82.45 | |
| DeepSeek-R1-Distill-Qwen-1.5BModel Scale=1.5B2026.01 | 82.1 | |
| GRESOBase Model=Qwen3-4B-Base, RL Method=GRESO, Rollout Configuration=Standard2026.01 | 80.8 | |
| DAPO_extendedBase Model=Qwen3-4B-Base, RL Method=DAPO, Rollout Configuration=2x Rollouts2026.01 | 80.8 | |
| GRESO_extendedBase Model=Qwen3-4B-Base, RL Method=GRESO, Rollout Configuration=2x Rollouts2026.01 | 80.6 | |
| DAPOBase Model=Qwen3-4B-Base, RL Method=DAPO, Rollout Configuration=Standard2026.01 | 80.2 | |
| Qwen2.5-Math-7B-InstructModel Scale=7B2026.01 | 79.81 | |
| GRPOBase Model=Qwen3-4B-Base, RL Method=GRPO, Rollout Configuration=Standard2026.01 | 79.4 | |
| Eurus-2-7B-PrimeModel Scale=7B2026.01 | 79.21 | |
| Rstar-Math-7BModel Scale=7B2026.01 | 78.4 | |
| GRPO_extendedBase Model=Qwen3-4B-Base, RL Method=GRPO, Rollout Configuration=2x Rollouts2026.01 | 78.2 | |
| Qwen3-8B-BaseBase Model=Qwen3-8B-Base, RL Method=None, Rollout Configuration=None2026.01 | 77 | |
| MATRIX-Gen-SFTBase model=Qwen-2.5-7B2024.10 | 73.6 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, RL Method=None, Rollout Configuration=None2026.01 | 72.2 | |
| Phi-4-mini + Mistral3-3B + OracleTraining=SD-E²2026.01 | 71.85 | |
| Tulu v2 MixBase model=Qwen-2.5-7B2024.10 | 65.2 | |
| WildChatBase model=Qwen-2.5-7B2024.10 | 63.4 | |
| OpenHermesBase model=Qwen-2.5-7B2024.10 | 59.8 | |
| TrajFusionBase Model=DeepSeekMath-7B, Train Size=100k2026.02 | 59.1 | |
| KPMath-PlusBase=Qwen1.5, Size=72B, Zero-shot without demonstrations=true2024.03 | 58.3 | |
| Evol-InstructBase model=Qwen-2.5-7B2024.10 | 57 | |
| Vanilla RFTBase Model=DeepSeekMath-7B, Train Size=100k2026.02 | 56.2 | |
| TrajFusionBase Model=DeepSeekMath-7B, Train Size=15k2026.02 | 55.8 | |
| LENS_GRPOBase Model=Llama3.2-3B-Instruct, RL Method=LENS, Rollout Configuration=Standard2026.01 | 55.8 | |
| ShareGPTBase model=Qwen-2.5-7B2024.10 | 53.9 | |
| DARTMathBase Model=DeepSeekMath-7B, Train Size=590k2026.02 | 53.6 | |
| MathFusionBase Model=DeepSeekMath-7B, Train Size=60k2026.02 | 53.4 | |
| GRESO_extendedBase Model=Llama3.2-3B-Instruct, RL Method=GRESO, Rollout Configuration=2x Rollouts2026.01 | 53.2 | |
| DAPOBase Model=Llama3.2-3B-Instruct, RL Method=DAPO, Rollout Configuration=Standard2026.01 | 53 | |
| DAPO_extendedBase Model=Llama3.2-3B-Instruct, RL Method=DAPO, Rollout Configuration=2x Rollouts2026.01 | 52.2 | |
| Mistral 24b-InstructType=Single-Agent Baseline, Search=Single-Pass2025.10 | 51.9 | |
| Vanilla RFTBase Model=DeepSeekMath-7B, Train Size=15k2026.02 | 51.8 | |
| GRESOBase Model=Llama3.2-3B-Instruct, RL Method=GRESO, Rollout Configuration=Standard2026.01 | 51.8 | |
| MCTS(10,3)+MASPRMAgent=Qwen2.5-1.5B-Instruct, Search=MCTS(10,3), Scorer=MASPRM-7B2025.10 | 51.8 | |
| GRPOBase Model=Llama3.2-3B-Instruct, RL Method=GRPO, Rollout Configuration=Standard2026.01 | 51.6 | |
| TrajFusionBase Model=LLaMA3-8B, Train Size=100k2026.02 | 51.4 | |
| SBS(5,3)+MASPRMAgent=Qwen2.5-1.5B-Instruct, Search=SBS(5,3), Scorer=MASPRM-7B2025.10 | 51.3 | |
| GRPO_extendedBase Model=Llama3.2-3B-Instruct, RL Method=GRPO, Rollout Configuration=2x Rollouts2026.01 | 51.2 | |
| GPT-4o-MiniType=Single-Agent Baseline, Search=Single-Pass2025.10 | 51 | |
| LEMMABase Model=DeepSeekMath-7B, Train Size=89k2026.02 | 50.6 | |
| KPMath-PlusBase=DSMath, Size=7B, Zero-shot without demonstrations=true2024.03 | 48.8 | |
| KPMath-PlusBase=Llemma, Size=34B, Zero-shot without demonstrations=true2024.03 | 48.6 | |
| KPMath-PlusBase=Llama-2, Size=70B, Zero-shot without demonstrations=true2024.03 | 48.6 | |
| Vanilla RFTBase Model=LLaMA3-8B, Train Size=100k2026.02 | 48.5 | |
| Ministral-3-8B-ReasoningTraining=SD-E²2026.01 | 47 | |
| InstructBase Model=DeepSeekMath-7B, Train Size=780k2026.02 | 46.9 | |
| KPMath-PlusBase=Mistral, Size=7B, Zero-shot without demonstrations=true2024.03 | 46.8 | |
| DARTMathBase Model=LLaMA3-8B, Train Size=590k2026.02 | 46.6 | |
| MathFusionBase Model=LLaMA3-8B, Train Size=60k2026.02 | 46.5 | |
| MMIQCBase Model=DeepSeekMath-7B, Train Size=2.3M2026.02 | 45.3 | |
| Qwen2.5-7B-InstructTraining=SD-E²2026.01 | 44 | |
| Magpie-SFTBase model=Qwen-2.5-7B2024.10 | 43.6 | |
| GPT-4 (0613)Zero-shot without demonstrations=false2024.03 | 42.5 | |
| KPMath-PlusBase=Llama-2, Size=13B, Zero-shot without demonstrations=true2024.03 | 41 | |
| Llama3.2-3B-InstructBase Model=Llama3.2-3B-Instruct, RL Method=None, Rollout Configuration=None2026.01 | 40.02 | |
| Tulu 3-SFTBase model=Qwen-2.5-7B2024.10 | 39.6 | |
| MMIQCBase Model=LLaMA3-8B, Train Size=2.3M2026.02 | 39.5 | |
| LEMMABase Model=LLaMA3-8B, Train Size=89k2026.02 | 38.3 | |
| TrajFusionBase Model=LLaMA3-8B, Train Size=15k2026.02 | 36.6 | |
| ChatGPTZero-shot without demonstrations=false2024.03 | 35.5 | |
| ICLBase Model=DeepSeekMath-7B2026.02 | 35.5 | |
| PaLM-2Size=540B, Zero-shot without demonstrations=false2024.03 | 34.3 | |
| Phi-3-small-8k-InstructTraining=SD-E²2026.01 | 34 | |
| Claude-2Zero-shot without demonstrations=false2024.03 | 32.5 | |
| MetaMathBase Model=LLaMA3-8B, Train Size=400k2026.02 | 32.5 | |
| MetaMath-Mistral-ProInference Protocol=Program of Thought (PoT)2024.01 | 30.3 | |
| MetaMath-Llemma-7BInference Protocol=Program of Thought (PoT)2024.01 | 30 | |
| Vanilla RFTBase Model=LLaMA3-8B, Train Size=15k2026.02 | 29.5 | |
| MetaMath-Mistral-7BInference Protocol=Program of Thought (PoT)2024.01 | 28.2 | |
| Llama3.1 8B-InstructType=Single-Agent Baseline, Search=Single-Pass2025.10 | 26.7 | |
| Llama3.2 3B-InstructType=Single-Agent Baseline, Search=Single-Pass2025.10 | 24.4 | |
| MetaMath-13BInference Protocol=Program of Thought (PoT)2024.01 | 22.4 | |
| DLRMain model=Frozen LLaMA-2-7B, #Params=1.3B+proj, #Forward=18%2026.01 | 22.1 | |
| ICLBase Model=LLaMA3-8B2026.02 | 21.2 | |
| WildChatBase model=Meta-Llama-3-8B2024.10 | 20.3 | |
| MetaMath-7BInference Protocol=Program of Thought (PoT)2024.01 | 19.8 | |
| MATRIX-Gen-SFTBase model=Meta-Llama-3-8B2024.10 | 19.3 | |
| Magpie-SFTBase model=Meta-Llama-3-8B2024.10 | 19.1 | |
| DeepSeekMath-RLMain model=LLaMA-2-7B, #Params=7B, #Forward=100%2026.01 | 18.3 | |
| Qwen2.5-1.5B-InstructType=Single-Agent Baseline, Search=Single-Pass2025.10 | 17.5 | |
| OpenHermesBase model=Meta-Llama-3-8B2024.10 | 16.8 | |
| GRPO (token-level)Main model=LLaMA-2-7B, #Params=7B, #Forward=100%2026.01 | 15.7 | |
| Tulu 3-SFTBase model=Meta-Llama-3-8B2024.10 | 15.6 |