Mathematical Reasoning on Minerva Math
63.2AccuracyPass@8 (Upper Bound)
Evaluation Results
| Method | Links | |
|---|---|---|
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-8, Bound=Upper Bound2026.01 | 63.2 | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-7B-Instruct2026.01 | 57.7 | |
| Qwen2.5-Math-PRM-7BTraining Samples=1500K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 54 | |
| Majority Vote@8Policy Model=Qwen2.5-14B-Instruct, Strategy=Majority Voting, Samples=82026.01 | 53.3 | |
| Qwen2.5-Math-7B-NAITTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 52.7 | |
| EurusPRM-Stage1Training Samples=463K, Aggregation Method=Min, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 52.6 | |
| EurusPRM-Stage2Training Samples=693K, Aggregation Method=Sum, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 51.5 | |
| Skywork-PRM-Qwen2.5-7BTraining Samples=N/A, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 51.1 | |
| Qwen2.5-Math-7B-MCRDTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 50.8 | |
| RLHFlow-PRM-Mistral-8BTraining Samples=273K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 50.4 | |
| RLHFlow-PRM-DeepSeek-8BTraining Samples=253K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 50 | |
| Math-Shepherd-PRM-7BTraining Samples=445K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 49.6 | |
| Qwen2.5-Math-7B-MCTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 49.3 | |
| SkyWork-OR1-Math-7BTeacher Model=SkyWork-OR1-Math-7B2026.03 | 49.3 | |
| GreedyPolicy Model=Qwen2.5-14B-Instruct, Strategy=Greedy Search2026.01 | 49.2 | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=-, Aggregation Method=Mean2026.01 | 48.5 | |
| SkyWork-OR1-7BTeacher Model=SkyWork-OR1-7B2026.03 | 47.1 | |
| Pass@8 (Upper Bound)Decoding Strategy=Pass@82025.05 | 46.7 | |
| Majority Vote@8Policy Model=Qwen2.5-7B-Instruct2026.01 | 46.7 | |
| Qwen2.5-Math-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=1500K, Aggregation Method=Mean2026.01 | 46.7 | |
| Qwen3-235B-A22BModel Category=Large Language Models, Model Scale=235B-A22B, Reasoning Strategy=Thinking2025.12 | 46.69 | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Pass@82026.01 | 46 | |
| Qwen2.5-Math-7B-NAITPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 45.8 | |
| EurusPRM-Stage2Policy Model=Qwen2.5-7B-Instruct, Training Samples=693K, Aggregation Method=Sum2026.01 | 45.6 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Thinking2025.12 | 45.22 | |
| Qwen2.5-Math-7B-MCRDPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 45.1 | |
| Math-Shepherd-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=445K, Aggregation Method=Mean2026.01 | 44.8 | |
| RLHFlow-PRM-DeepSeek-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=253K, Aggregation Method=Mean2026.01 | 44.8 | |
| EurusPRM-Stage1Policy Model=Qwen2.5-7B-Instruct, Training Samples=463K, Aggregation Method=Min-Max2026.01 | 44.8 | |
| FIGR2025.12 | 44.49 | |
| Qwen2.5-Math-7B-MCPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 44.2 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 43.77 | |
| Text-only RLMode=Text-only, Reasoning Strategy=RL2025.12 | 43.38 | |
| GreedyPolicy Model=Qwen2.5-7B-Instruct2026.01 | 43 | |
| R1-Qwen-1.5B + EvoCoTBackbone=R1-Qwen-1.5B, Algorithm=EvoCoT2025.08 | 42.8 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 42.67 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 42.57 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 42.53 | |
| RLHFlow-PRM-Mistral-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=273K, Aggregation Method=Mean2026.01 | 42.3 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 42.2 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 41.37 | |
| Qwen3-VL-32B-InstructModel Category=Large Vision-Language Models, Model Scale=32B, Reasoning Strategy=Instruct2025.12 | 41.18 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 40.91 | |
| Majority @8Decoding Strategy=Majority @82025.05 | 40.1 | |
| Qwen3-VL-8B-InstructModel Category=Large Vision-Language Models, Model Scale=8B, Reasoning Strategy=Instruct2025.12 | 40.07 | |
| Hi-CoTModel=DeepSeek-R1-Distill-Qwen-32B2026.03 | 39.7 | |
| Qwen2.5-7B + PRIME (380K)Backbone=Qwen2.5-7B, Algorithm=PRIME, Training Data Size=380K2025.08 | 39.7 | |
| RLOOBackbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 39.48 | |
| GRPOBackbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 39.39 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Non-Thinking2025.12 | 39.34 | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 4B, Transfer Setting=Second Setting (Red Shading)2026.04 | 39.2 | |
| Skywork-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 39.1 | |
| GRPOBackbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 39.07 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 38.75 | |
| Majority Vote@8Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Majority Vote@82026.01 | 38.6 | |
| GLM-4.5VModel Category=Large Vision-Language Models, Model Scale=108B2025.12 | 38.6 | |
| REOPOLDTeacher Model=SkyWork-OR1-Math-7B2026.03 | 38.6 | |
| Hi-CoT (format-relaxed)Model=DeepSeek-R1-Distill-Qwen-32B2026.03 | 38.6 | |
| RLHFlow-PRM-Deepseek-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 38.2 | |
| Skywork-PRM-1.5BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 38.2 | |
| SCOPEDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 38.2 | |
| SCOPE (w/o AST normalization)Decoding Strategy=Best-of-8 strategy, Ablation=w/o AST normalization2025.05 | 38.2 | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Aggregation Method=Mean2026.01 | 38.2 | |
| R1-Qwen-1.5B + DeepScaleR(GRPO)Backbone=R1-Qwen-1.5B, Algorithm=DeepScaleR(GRPO)2025.08 | 38.2 | |
| Qwen2.5-Math-7B-PRM800KDecoding Strategy=Best-of-8 strategy, Training Data=manually annotated (PRM800K)2025.05 | 38.1 | |
| EurusPRM-Stage1Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 37.9 | |
| PRISMModel=Qwen2.5-7B, Data=DAPO-17k2026.01 | 37.9 | |
| EurusPRM-Stage2Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 37.7 | |
| RLHFlow-PRM-Mistral-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 37.5 | |
| Qwen2.5-Math-PRM-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=1500K, Aggregation Method=Mean2026.01 | 37.5 | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 3B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 37.4 | |
| Token-ALPevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 37.27 | |
| SCOPE (Step Replacement)Decoding Strategy=Best-of-8 strategy, Ablation=Step Replacement2025.05 | 37.1 | |
| SCOPE (Step Skipping)Decoding Strategy=Best-of-8 strategy, Ablation=Step Skipping2025.05 | 37.1 | |
| Qwen2.5-Math-7B-NAITPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=128K, Aggregation Method=Mean2026.01 | 37.1 | |
| Hi-CoTModel=Qwen3-32B2026.03 | 37.1 | |
| Qwen2.5-7B + EvoCoTBackbone=Qwen2.5-7B, Algorithm=EvoCoT2025.08 | 37.1 | |
| Seq-ALPevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 37.06 | |
| LLaDA RCDModel=LLaDA, Sequence length=512, Decoding strategy=single-token-per-step2026.01 | 37 | |
| INT.Model=Qwen2.5-7B, Data=DAPO-17k2026.01 | 36.8 | |
| Seq-MISevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 36.75 | |
| Qwen2.5-Math-7B-MCRDPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=128K, Aggregation Method=Mean2026.01 | 36.7 | |
| GRPOevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 36.43 | |
| SCOPE (w/o code translation)Decoding Strategy=Best-of-8 strategy, Ablation=w/o code translation2025.05 | 36.4 | |
| RLHFlow-PRM-Mistral-8BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=273K, Aggregation Method=Mean2026.01 | 36.4 | |
| EurusPRM-Stage1Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=463K, Aggregation Method=Min-Max2026.01 | 36.4 | |
| SFTTeacher Model=SkyWork-OR1-Math-7B2026.03 | 36.4 | |
| RLOOBackbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 36.07 | |
| Hi-CoTModel=Qwen3-4B-Instruct-25072026.03 | 36 | |
| StandardModel=DeepSeek-R1-Distill-Qwen-14B2026.03 | 36 | |
| Token-MISevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 35.94 | |
| Hi-CoT (format-relaxed)Model=Qwen3-4B-Thinking-25072026.03 | 35.7 | |
| Hi-CoT (format-relaxed)Model=Qwen3-32B2026.03 | 35.7 | |
| Hi-CoTModel=DeepSeek-R1-Distill-Qwen-14B2026.03 | 35.7 | |
| EurusPRM-Stage2Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=693K, Aggregation Method=Sum2026.01 | 35.6 | |
| CoTModel=Qwen3-4B-Instruct-25072026.03 | 35.3 | |
| Hi-CoTModel=Qwen3-4B-Thinking-25072026.03 | 35.3 | |
| Hi-CoT (format-relaxed)Model=DeepSeek-R1-Distill-Qwen-14B2026.03 | 35.3 | |
| Seq-Bypassevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 35.23 | |
| RLHFlow-PRM-DeepSeek-8BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=253K, Aggregation Method=Mean2026.01 | 34.9 |