Mathematical Reasoning on Gaokao
80AccuracyPass@8 (Upper Bound)
Evaluation Results
| Method | Links | |
|---|---|---|
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-8, Bound=Upper Bound2026.01 | 80 | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-7B-Instruct2026.01 | 77.4 | |
| Skywork-PRM-Qwen2.5-7BTraining Samples=N/A, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 74.5 | |
| Qwen2.5-Math-PRM-7BTraining Samples=1500K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 73.5 | |
| Majority Vote@8Policy Model=Qwen2.5-14B-Instruct, Strategy=Majority Voting, Samples=82026.01 | 73.2 | |
| Qwen2.5-Math-7B-NAITTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 73.2 | |
| Qwen2.5-Math-7B-MCRDTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 72.9 | |
| RLHFlow-PRM-DeepSeek-8BTraining Samples=253K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 72.7 | |
| EurusPRM-Stage2Training Samples=693K, Aggregation Method=Sum, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 72.2 | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=-, Aggregation Method=Mean2026.01 | 72.2 | |
| DPO (Random)Base Model=Qwen2.5-7B-Instruct2026.02 | 71.4 | |
| SAGEBase Model=Qwen2.5-7B-Instruct2026.02 | 71.4 | |
| EurusPRM-Stage2Policy Model=Qwen2.5-7B-Instruct, Training Samples=693K, Aggregation Method=Sum2026.01 | 71.4 | |
| Qwen2.5-Math-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=1500K, Aggregation Method=Mean2026.01 | 71.4 | |
| Qwen2.5-Math-7B-NAITPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 71.3 | |
| EurusPRM-Stage1Training Samples=463K, Aggregation Method=Min, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 71.2 | |
| RLHFlow-PRM-Mistral-8BTraining Samples=273K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 70.9 | |
| DPO (Full)Base Model=Qwen2.5-7B-Instruct2026.02 | 70.6 | |
| EurusPRM-Stage1Policy Model=Qwen2.5-7B-Instruct, Training Samples=463K, Aggregation Method=Min-Max2026.01 | 70.3 | |
| Majority Vote@8Policy Model=Qwen2.5-7B-Instruct2026.01 | 70.1 | |
| VanillaBase Model=Qwen2.5-7B-Instruct2026.02 | 69.9 | |
| Math-Shepherd-PRM-7BTraining Samples=445K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 69.6 | |
| Pass@8 (Upper Bound)Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Pass@82026.01 | 69.6 | |
| GreedyPolicy Model=Qwen2.5-14B-Instruct, Strategy=Greedy Search2026.01 | 69.3 | |
| Qwen2.5-Math-7B-MCTraining Samples=128K, Aggregation Method=Mean, Policy Model=Qwen2.5-14B-Instruct, Strategy=Best-of-82026.01 | 69.1 | |
| Qwen2.5-Math-7B-MCRDPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 68.6 | |
| RLHFlow-PRM-DeepSeek-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=253K, Aggregation Method=Mean2026.01 | 68.3 | |
| GreedyPolicy Model=Qwen2.5-7B-Instruct2026.01 | 67 | |
| Math-Shepherd-PRM-7BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=445K, Aggregation Method=Mean2026.01 | 67 | |
| RLHFlow-PRM-Mistral-8BPolicy Model=Qwen2.5-7B-Instruct, Training Samples=273K, Aggregation Method=Mean2026.01 | 67 | |
| Qwen2.5-Math-7B-MCPolicy Model=Qwen2.5-7B-Instruct, Training Samples=128K, Aggregation Method=Mean2026.01 | 66.8 | |
| Majority Vote@8Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Majority Vote@82026.01 | 66.5 | |
| Qwen2.5-Math-PRM-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=1500K, Aggregation Method=Mean2026.01 | 63.6 | |
| Skywork-PRM-Qwen2.5-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Aggregation Method=Mean2026.01 | 63.1 | |
| Qwen2.5-Math-7B-NAITPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=128K, Aggregation Method=Mean2026.01 | 62.8 | |
| EurusPRM-Stage2Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=693K, Aggregation Method=Sum2026.01 | 60.8 | |
| EurusPRM-Stage1Policy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=463K, Aggregation Method=Min-Max2026.01 | 60.2 | |
| GreedyPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Greedy2026.01 | 60 | |
| Qwen2.5-Math-7B-MCRDPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=128K, Aggregation Method=Mean2026.01 | 59.9 | |
| RLHFlow-PRM-Mistral-8BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=273K, Aggregation Method=Mean2026.01 | 59.5 | |
| RLHFlow-PRM-DeepSeek-8BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=253K, Aggregation Method=Mean2026.01 | 59.2 | |
| Math-Shepherd-PRM-7BPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=445K, Aggregation Method=Mean2026.01 | 58.7 | |
| Qwen2.5-Math-7B-MCPolicy Model=Qwen2.5-3B-Instruct, Evaluation Protocol=Best-of-8, Training Samples=128K, Aggregation Method=Mean2026.01 | 58.4 | |
| SAGEBase Model=Qwen2.5-3B-Instruct2026.02 | 58.23 | |
| DPO (Full)Base Model=Qwen2.5-3B-Instruct2026.02 | 56.9 | |
| VanillaBase Model=Qwen2.5-3B-Instruct2026.02 | 56.4 | |
| DPO (Random)Base Model=Qwen2.5-3B-Instruct2026.02 | 56.4 | |
| SAGEBase Model=Qwen2.5-1.5B-Instruct2026.02 | 50.4 | |
| DPO (Random)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 48.6 | |
| DPO (Full)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 47.3 | |
| VanillaBase Model=Qwen2.5-1.5B-Instruct2026.02 | 46.2 |