Mathematical Reasoning on Minerva
55.96Pass@1MulFeRL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MulFeRLBackbone=Qwen3-4B-Inst, Training Strategy=Multi-turn Feedback-guided Reinforcement Learning2026.01 | 55.96 | — | — | — | — | — | |
| GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 51.25 | — | — | — | — | — | |
| Critique-GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 50.29 | — | — | — | — | — | |
| Dr.GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 49.19 | — | — | — | — | — | |
| CITL-FTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 47.13 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-14B + iGRPOBase Model=DeepSeek-R1-Distill-Qwen-14B, Training Method=iGRPO, Parameter Size=14B2026.02 | 47.06 | — | — | — | — | — | |
| SFTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 46.32 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-14B + GRPOBase Model=DeepSeek-R1-Distill-Qwen-14B, Training Method=GRPO, Parameter Size=14B2026.02 | 46 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-14BBase Model=DeepSeek-R1-Distill-Qwen-14B, Parameter Size=14B2026.02 | 45.59 | — | — | — | — | — | |
| APOBackbone=Qwen2.5-Math-7B, K=162026.02 | 45.5 | 66.37 | — | — | — | — | |
| RAFTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 45.29 | — | — | — | — | — | |
| GRPO (Baseline)Backbone=Qwen2.5-Math-7B, K=162026.02 | 45.01 | 65.44 | — | — | — | — | |
| NSRBackbone=Qwen2.5-Math-7B, K=162026.02 | 44.93 | 66.12 | — | — | — | — | |
| Qwen3-4B-InstBackbone=Qwen3-4B-Inst, Training Strategy=Base Model2026.01 | 42.43 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7B + iGRPOBase Model=DeepSeek-R1-Distill-Qwen-7B, Training Method=iGRPO, Parameter Size=7B2026.02 | 41.54 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7B + Critique-GRPOBase Model=DeepSeek-R1-Distill-Qwen-7B, Training Method=Critique-GRPO, Parameter Size=7B2026.02 | 41.1 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7B + Self-VerificationBase Model=DeepSeek-R1-Distill-Qwen-7B, Training Method=Self-Verification, Parameter Size=7B2026.02 | 41 | — | — | — | — | — | |
| GRPO (Baseline)Backbone=Qwen2.5-7B, K=162026.02 | 40.65 | 63.97 | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7B + GRPOBase Model=DeepSeek-R1-Distill-Qwen-7B, Training Method=GRPO, Parameter Size=7B2026.02 | 40.44 | — | — | — | — | — | |
| SimpleRL-7BModel Scale=7B2026.01 | 39.72 | — | — | — | — | — | |
| R³-7BModel Scale=7B2026.01 | 39.7 | — | — | — | — | — | |
| APOBackbone=Qwen2.5-7B, K=162026.02 | 39.17 | 66.91 | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Parameter Size=7B2026.02 | 39.1 | — | — | — | — | — | |
| DoPRBackbone=Qwen, Model Size=7B2026.01 | 39 | — | — | — | — | — | |
| SATURN-7BModel Scale=7B2026.01 | 38.96 | — | — | — | — | — | |
| UFOBackbone=Qwen, Model Size=7B2026.01 | 38.7 | — | — | — | — | — | |
| Eurus-2-7B-PrimeModel Scale=7B2026.01 | 38.62 | — | — | — | — | — | |
| Multiplex ThinkingBackbone=DeepSeek-R1-Distill-Qwen-7B2026.01 | 38.6 | — | — | — | — | — | |
| OpenMath-Nemotron-14B + iGRPOBackbone=Nemotron-14B, Fine-tuning=iGRPO2026.02 | 38.24 | — | — | — | — | — | |
| Thinker-7BModel Scale=7B2026.01 | 38.16 | — | — | — | — | — | |
| MulFeRLBackbone=Qwen2.5-7B-Base, Training Strategy=Multi-turn Feedback-guided Reinforcement Learning2026.01 | 37.87 | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B2026.01 | 37.71 | — | — | — | — | — | |
| M2POModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 37.7 | — | — | — | — | — | |
| GRPO-KLBackbone=Qwen2.5-Math-7B, K=162026.02 | 37.68 | 67.28 | — | — | — | — | |
| s1.1-7BBackbone=Qwen2.5-7B, Pure SFT?=false2025.12 | 37.5 | — | — | — | — | — | |
| Stochastic Soft ThinkingBackbone=DeepSeek-R1-Distill-Qwen-7B2026.01 | 37.2 | — | — | — | — | — | |
| GRPO-KL (Error-Only)Backbone=Qwen2.5-Math-7B, K=162026.02 | 37.11 | 66.18 | — | — | — | — | |
| OpenMath-Nemotron-14B + iGRPOBase Model=OpenMath-Nemotron-14B, Training Method=iGRPO, Parameter Size=14B2026.02 | 36.76 | — | — | — | — | — | |
| Qwen InstructBackbone=Qwen2.5-7B, Pure SFT?=false2025.12 | 35.8 | — | — | — | — | — | |
| OpenMath-Nemotron-14B + GRPOBase Model=OpenMath-Nemotron-14B, Training Method=GRPO, Parameter Size=14B2026.02 | 35.7 | — | — | — | — | — | |
| CISPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 35.7 | — | — | — | — | — | |
| Discrete RLBackbone=DeepSeek-R1-Distill-Qwen-7B2026.01 | 35.3 | — | — | — | — | — | |
| LLaDA1.5Corrector Sampling=false, Number of shots=42026.02 | 35.1 | — | — | — | — | — | |
| ProSeCo SamplingCorrector Sampling=true, Number of shots=42026.02 | 35.1 | — | — | — | — | — | |
| OpenMath-Nemotron-7B + iGRPOBase Model=OpenMath-Nemotron-7B, Training Method=iGRPO, Parameter Size=7B2026.02 | 34.94 | — | — | — | — | — | |
| Qwen2.5-Math-7B-InstructModel Scale=7B2026.01 | 34.6 | — | — | — | — | — | |
| R³-1.5BModel Scale=1.5B2026.01 | 34.21 | — | — | — | — | — | |
| OpenMath-Nemotron-7B + GRPOBase Model=OpenMath-Nemotron-7B, Training Method=GRPO, Parameter Size=7B2026.02 | 34 | — | — | — | — | — | |
| Critique-GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 33.82 | — | — | — | — | — | |
| M2POModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 33.8 | — | — | — | — | — | |
| MinPROModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 33.6 | — | — | — | — | — | |
| GSPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 33.5 | — | — | — | — | — | |
| QLoRABackbone=Qwen2.5-7B, Pure SFT?=true2025.12 | 33.5 | — | — | — | — | — | |
| OpenMath-Nemotron-7BBase Model=OpenMath-Nemotron-7B, Parameter Size=7B2026.02 | 33.46 | — | — | — | — | — | |
| OpenMath-Nemotron-14BBase Model=OpenMath-Nemotron-14B, Parameter Size=14B2026.02 | 33.46 | — | — | — | — | — | |
| OpenMath-Nemotron-14BBackbone=Nemotron-14B2026.02 | 33.46 | — | — | — | — | — | |
| MathForgeBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 33.36 | — | — | — | — | — | |
| LLaDA-InstructCorrector Sampling=false, Number of shots=42026.02 | 33.32 | — | — | — | — | — | |
| MinPROModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 33.3 | — | — | — | — | — | |
| Discrete CoTBackbone=DeepSeek-R1-Distill-Qwen-7B2026.01 | 33.3 | — | — | — | — | — | |
| GRPO-KLBackbone=Qwen2.5-7B, K=162026.02 | 32.97 | 64.71 | — | — | — | — | |
| FASTCURL-1.5B-V2Model Scale=1.5B2026.01 | 32.81 | — | — | — | — | — | |
| Nemotron-H-8B-Base-8K + iGRPOBase Model=Nemotron-H-8B-Base-8K, Training Method=iGRPO, Parameter Size=8B2026.02 | 32.72 | — | — | — | — | — | |
| LLaDA-Instruct + ReMDMCorrector Sampling=true, Number of shots=42026.02 | 32.72 | — | — | — | — | — | |
| ProSeCo SFTCorrector Sampling=false, Number of shots=42026.02 | 32.42 | — | — | — | — | — | |
| CISPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 32.2 | — | — | — | — | — | |
| RePOTraining Set=OpenR1-Math Hard Subset2026.02 | 32 | — | — | — | — | — | |
| Qwen InstructBackbone=Qwen2.5-Math-1.5B, Pure SFT?=false2025.12 | 31.6 | — | — | — | — | — | |
| Nemotron-H-8B-Base-8K + Critique-GRPOBase Model=Nemotron-H-8B-Base-8K, Training Method=Critique-GRPO, Parameter Size=8B2026.02 | 31.5 | — | — | — | — | — | |
| MQRBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 31.43 | — | — | — | — | — | |
| GSPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 31.4 | — | — | — | — | — | |
| Nemotron-H-8B-Base-8K + Self-VerificationBase Model=Nemotron-H-8B-Base-8K, Training Method=Self-Verification, Parameter Size=8B2026.02 | 31.1 | — | — | — | — | — | |
| GRPOModel Scale=14B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 31.1 | — | — | — | — | — | |
| Llama3.1-InstructCorrector Sampling=false, Number of shots=42026.02 | 31.1 | — | — | — | — | — | |
| DGPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 31.07 | — | — | — | — | — | |
| Dr.GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 30.37 | — | — | — | — | — | |
| DeepScaleR-1.5B-PreviewModel Scale=1.5B2026.01 | 30.35 | — | — | — | — | — | |
| Vanilla SFT + ReMDMCorrector Sampling=true, Training strategy=SFT, Number of shots=42026.02 | 29.9 | — | — | — | — | — | |
| GRPOBackbone=Qwen, Model Size=1.5B2026.01 | 29.8 | — | — | — | — | — | |
| Vanilla SFTCorrector Sampling=false, Training strategy=SFT, Number of shots=42026.02 | 29.74 | — | — | — | — | — | |
| UFOBackbone=Qwen, Model Size=1.5B2026.01 | 29.7 | — | — | — | — | — | |
| Nemotron-H-8B-Base-8K + GRPOBase Model=Nemotron-H-8B-Base-8K, Training Method=GRPO, Parameter Size=8B2026.02 | 29.56 | — | — | — | — | — | |
| DAPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 29.5 | — | — | — | — | — | |
| GRPOBackbone=Qwen, Model Size=7B2026.01 | 29.4 | — | — | — | — | — | |
| DoPRBackbone=Qwen, Model Size=1.5B2026.01 | 29.4 | — | — | — | — | — | |
| LadderBackbone=Qwen2.5-Math-1.5B, Pure SFT?=true2025.12 | 29.4 | — | — | — | — | — | |
| LadderBackbone=Qwen2.5-7B, Pure SFT?=true2025.12 | 29.3 | — | — | — | — | — | |
| GRPO-ADBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 29.14 | — | — | — | — | — | |
| QLoRABackbone=Qwen2.5-Math-1.5B, Pure SFT?=true2025.12 | 29.1 | — | — | — | — | — | |
| GRPOModel Scale=8B-Base, Off-policyness=Large, Generations per prompt=22026.01 | 29 | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 28.97 | — | — | — | — | — | |
| STILL-3-1.5BModel Scale=1.5B2026.01 | 28.81 | — | — | — | — | — | |
| Dr.GRPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 28.58 | — | — | — | — | — | |
| GSPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 28.12 | — | — | — | — | — | |
| TAMPOMaximum response length=6k tokens, Training algorithm=TAMPO2026.02 | 27.9 | — | — | — | 44.8 | — | |
| LLaDA-BaseCorrector Sampling=false, Number of shots=42026.02 | 27.88 | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 27.76 | — | — | — | — | — | |
| GPGBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 27.21 | — | — | — | — | — | |
| Multiplex ThinkingBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 26.2 | — | — | — | — | — | |
| GRPO (Ts : 0.9)Sampling temperature (Ts)=0.9, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 26.1 | — | — | — | 43.4 | — |