Mathematical Reasoning on Minerva (pass@1 accuracy %)
74.8Pass@1 AccuracyQwen2.5-32B + SimpleRL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-32B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=32B, Training Strategy=SimpleRL2026.04 | 74.8 | |
| Qwen2.5-14B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=14B, Training Strategy=SimpleRL2026.04 | 74.3 | |
| Qwen2.5-7B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Training Strategy=SimpleRL2026.04 | 65.7 | |
| Qwen2.5-7B + SimpleRL (Lighteval)Evaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Training Strategy=SimpleRL, Evaluation Framework=Lighteval2026.04 | 65.3 | |
| Qwen2.5-32BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=32B2026.04 | 56 | |
| Qwen3Model Category=Original Models2026.05 | 52.9 | |
| HiPO-Mt-biasBackbone=Qwen2.5-7B-Instruct2026.04 | 51.02 | |
| TEPOBackbone=Qwen3-14B, Method Variant=TEPO2026.04 | 49.26 | |
| DPOBackbone=Qwen2.5-7B-Instruct2026.04 | 48.81 | |
| Qwen3-14B w. Entropy-based TermBackbone=Qwen3-14B, Method Variant=Entropy-based Term2026.04 | 48.52 | |
| Qwen3-14B w. GSPOBackbone=Qwen3-14B, Method Variant=GSPO2026.04 | 48.52 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.04 | 48.26 | |
| Qwen3-14B w. GRPO/DAPOBackbone=Qwen3-14B, Method Variant=GRPO/DAPO2026.04 | 47.79 | |
| Qwen3-14B w. KL-CovBackbone=Qwen3-14B, Method Variant=KL-Cov2026.04 | 47.79 | |
| HiPO-Rq-biasBackbone=Qwen2.5-7B-Instruct2026.04 | 47.66 | |
| HiPO-Rq+Mt-biasBackbone=Qwen2.5-7B-Instruct2026.04 | 47.65 | |
| Qwen2.5-14BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=14B2026.04 | 47.6 | |
| Qwen2.5-7BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=7B2026.04 | 47.4 | |
| Qwen3-8192Model Category=Original Models, Max Generation Length=81922026.05 | 47.4 | |
| Qwen3-14B w. GPGBackbone=Qwen3-14B, Method Variant=GPG2026.04 | 47.05 | |
| Qwen2.5-7B (Lighteval)Evaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Evaluation Framework=Lighteval2026.04 | 46.4 | |
| Qwen3-14BBackbone=Qwen3-14B, Method Variant=Base2026.04 | 45.95 | |
| Qwen3-14B w. CLIP-CovBackbone=Qwen3-14B, Method Variant=CLIP-Cov2026.04 | 45.58 | |
| DeepSeek-R1-Distill-32BModel Name=DeepSeek-R1-Distill-32B, Model Size=32B, Method Type=Baseline2026.04 | 45.1 | |
| Qwen2.5-7B-Math + SFT + RL-TCERBase Model=Qwen2.5-7B-Math, Training Strategy=SFT + RL-TCER2026.04 | 44.5 | |
| RL-PLUSBase Model=Qwen2.5-Math-7B, Training Method=RL-PLUS2025.07 | 43.8 | |
| RL-PLUSBase Model=Qwen2.5-Math-7B2025.07 | 43.8 | |
| Qwen2.5-32B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=32B, Training Strategy=SimpleRL2026.04 | 43.5 | |
| SCRLCandidate responses=64, Training samples=32, Backbone=Qwen2.5-Math-7B2026.03 | 43.1 | |
| Qwen2.5-7B-Math + SFT + RL-EndoRBase Model=Qwen2.5-7B-Math, Training Strategy=SFT + RL-EndoR2026.04 | 42.6 | |
| HölderPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 42.3 | |
| SFTModel Category=Off-Policy and Mixed-Policy2026.05 | 42.3 | |
| Qwen2.5-14B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=14B, Training Strategy=SimpleRL2026.04 | 41.7 | |
| SCRLCandidate responses=32, Training samples=16, Backbone=Qwen2.5-Math-7B2026.03 | 41.6 | |
| PieceHint-Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=PieceHint2026.04 | 41.6 | |
| Re-Schedule_sigmoidBackbone=Qwen2.5-7B, Weight Function=sigmoid2025.10 | 41.5 | |
| Re-Schedule_linearBackbone=Qwen2.5-7B, Weight Function=linear2025.10 | 41.2 | |
| Qwen2.5-7B-Math + SFTBase Model=Qwen2.5-7B-Math, Training Strategy=SFT2026.04 | 40.8 | |
| SFTBase Model=Qwen2.5-Math-7B, Training Method=SFT2025.07 | 40.8 | |
| SFTBase Model=Qwen2.5-Math-7B2025.07 | 40.8 | |
| Qwen3-4BModel Name=Qwen3-4B, Model Size=4B, Method Type=Baseline2026.04 | 40.4 | |
| TGPO-annealingModel Category=On-Policy Methods2026.05 | 40.4 | |
| DAPOBase Model=Qwen2.5-Math-7B2025.07 | 40.1 | |
| ReLIFTBase Model=Qwen2.5-Math-7B2025.07 | 40.1 | |
| DeepSeek-R1-Distill-7BModel Name=DeepSeek-R1-Distill-7B, Model Size=7B, Method Type=Baseline2026.04 | 40 | |
| SFT+GRPOBase Model=Qwen2.5-Math-7B2025.07 | 39.7 | |
| PRIMEFine-tuning base=Qwen2.5-Math-7B, Comparison Group=Non-comparable Baselines2025.08 | 39.7 | |
| GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, Hölder Parameter (p)=12026.05 | 39.7 | |
| GRPOBase Model=Qwen2.5-Math-7B, Training Method=GRPO2025.07 | 39.3 | |
| GRPOBase Model=Qwen2.5-Math-7B2025.07 | 39.3 | |
| PMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 39.3 | |
| PRIMEBase Model=Qwen2.5-Math-7B2025.07 | 39 | |
| PRIME-ZeroModel Category=On-Policy Methods2026.05 | 39 | |
| GRPOBackbone=Qwen2.5-7B, Category=Classical RLVR Methods2025.10 | 38.6 | |
| LUFFYModel Category=Off-Policy and Mixed-Policy2026.05 | 38.6 | |
| TAPOBase Model=Qwen2.5-Math-7B2025.07 | 38.2 | |
| GMPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=p → 02026.05 | 37.9 | |
| TGPOModel Category=On-Policy Methods2026.05 | 37.9 | |
| Qwen2.5-7B w. CLIP-CovBackbone=Qwen2.5-7B, Method Variant=CLIP-Cov2026.04 | 37.86 | |
| Eurus-2-7B-PRIMEZero-shot=true, Model Category=Qwen 7B Models, Training Samples=230K + 150K2025.05 | 37.5 | |
| LUFFYBase Model=Qwen2.5-Math-7B2025.07 | 37.5 | |
| HölderPOTraining Stage=RL Post-Trained Models, Backbone=Qwen3-8B-Base, Training Objective=Linear Des: 2 → -22026.05 | 37.5 | |
| Dr.GRPOModel Scale=7B, Base Architecture=R1-Distill-Qwen-7B, Training Phase/Objective=RL Post-Trained2026.05 | 37.5 | |
| GRPO++Model Category=On-Policy Methods2026.05 | 37.5 | |
| Self-Instruct + CoT-Self-InstructBackbone=Qwen3-8B-Base, # Train=63882025.11 | 37.27 | |
| InstructFine-tuning base=Qwen2.5-Math-7B, Comparison Group=Non-comparable Baselines2025.08 | 37.1 | |
| PieceHint-DeepSeek-1.5BModel Name=DeepSeek-R1-Distill-1.5B, Model Size=1.5B, Method Type=PieceHint2026.04 | 36.7 | |
| Qwen2.5-7B w. Entropy-based TermBackbone=Qwen2.5-7B, Method Variant=Entropy-based Term2026.04 | 36.02 | |
| GPG-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 36 | |
| DPO-LoRRFine-tuning base=Qwen2.5-Math-7B, Iteration=iter32025.08 | 36 | |
| PRIME-Zero-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 36 | |
| Eurus-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 36 | |
| KDRLModel Category=On-Policy Methods2026.05 | 36 | |
| Self-InstructBackbone=Qwen3-8B-Base, # Train=59902025.11 | 35.79 | |
| CoT Cold-Start + Solver FeedbackBackbone=Qwen3-8B-Base, # Train=68012025.11 | 35.79 | |
| Qwen2.5-7B w. GSPOBackbone=Qwen2.5-7B, Method Variant=GSPO2026.04 | 35.66 | |
| EGRSDBackbone=Qwen3-8B2026.05 | 35.48 | |
| Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=Baseline2026.04 | 35.4 | |
| AAPO-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 35.3 | |
| LPPOBackbone=Qwen2.5-7B, Category=Scheduling Methods2025.10 | 35.3 | |
| Qwen2.5-7B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=7B, Training Strategy=SimpleRL2026.04 | 35.1 | |
| OPSDBackbone=Qwen3-8B2026.05 | 34.94 | |
| Qwen2.5-7B w. KL-CovBackbone=Qwen2.5-7B, Method Variant=KL-Cov2026.04 | 34.92 | |
| TEPOBackbone=Qwen2.5-7B, Method Variant=TEPO2026.04 | 34.92 | |
| Oat-Zero-7BZero-shot=true, Model Category=Qwen 7B Models, Training Samples=8,5232025.05 | 34.9 | |
| GRPO w/ SFT LossBase Model=Qwen2.5-Math-7B2025.07 | 34.9 | |
| PMPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 34.9 | |
| HölderPOModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained, p-Scheduling Strategy=Linear Des: 2 → -22026.05 | 34.9 | |
| CL-EGRSDBackbone=Qwen3-8B2026.05 | 34.83 | |
| SCRLBase Model=Qwen3-14B-Base2026.05 | 34.7 | |
| GRPOBackbone=Qwen3-8B2026.05 | 34.65 | |
| OatBase Model=Qwen2.5-Math-7B2025.07 | 34.6 | |
| Oat-ZeroModel Category=On-Policy Methods2026.05 | 34.6 | |
| Self-Instruct + RLMTBackbone=Qwen3-8B-Base, # Train=71282025.11 | 34.56 | |
| Qwen2.5-7B w. GPGBackbone=Qwen2.5-7B, Method Variant=GPG2026.04 | 34.55 | |
| CRISPBackbone=Qwen3-8B2026.05 | 34.47 | |
| Baseline (no-train)Backbone=Qwen3-8B2026.05 | 34.32 | |
| ACC_sigmoidBackbone=Qwen2.5-7B, Category=Scheduling Methods2025.10 | 34.2 | |
| GPG-7BModel Scale=7B, Base Architecture=Qwen2.5-Math, Training Phase/Objective=RL Post-Trained2026.05 | 34.2 | |
| QuestABase Model=Qwen3-14B-Base2026.05 | 34 |