Mathematical Reasoning on GaoKao En 2023
79.9Pass@1 AccuracyCES
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CESBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 79.9 | 1,914 | |
| Pass@8 (Upper Bound)Decoding Strategy=Pass@82025.05 | 79.7 | — | |
| GPT-o1-miniDecoding=Greedy2025.02 | 78.4 | — | |
| Scaf-GRPOBase Model=DeepSeek-R1-Distill-Qwen-1.5B2025.10 | 72.3 | — | |
| SCOPE (w/o AST normalization)Decoding Strategy=Best-of-8 strategy, Ablation=w/o AST normalization2025.05 | 72.2 | — | |
| DAPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 72.2 | 2,182 | |
| SCOPEDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 71.9 | — | |
| R1-7BBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 71.8 | 2,511 | |
| Entropy ShapeBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 71.8 | 2,048 | |
| SCOPE (Step Skipping)Decoding Strategy=Best-of-8 strategy, Ablation=Step Skipping2025.05 | 71.7 | — | |
| Vanilla GRPOBase Model=DeepSeek-R1-Distill-Qwen-1.5B2025.10 | 71.4 | — | |
| SCOPE (Step Replacement)Decoding Strategy=Best-of-8 strategy, Ablation=Step Replacement2025.05 | 71 | — | |
| RLHFlow-PRM-Mistral-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 70.9 | — | |
| Skywork-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 70.9 | — | |
| Skywork-PRM-1.5BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 70.6 | — | |
| SCOPE (w/o code translation)Decoding Strategy=Best-of-8 strategy, Ablation=w/o code translation2025.05 | 70.6 | — | |
| Qwen2.5-Math-7B-PRM800KDecoding Strategy=Best-of-8 strategy, Training Data=manually annotated (PRM800K)2025.05 | 70.5 | — | |
| Majority @8Decoding Strategy=Majority @82025.05 | 70.4 | — | |
| Qwen2.5-Math-7B-S2R-ORLBackbone=Qwen2.5-Math-7B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 70.1 | — | |
| RLHFlow-PRM-Deepseek-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 69.9 | — | |
| OpenReasoner-Zero-7BModel=Qwen2.5-Math-7B, Method type=Reinforcement Learning, Result source=-2025.06 | 68.8 | — | |
| GPT-4oDecoding=Greedy2025.02 | 67.5 | — | |
| EurusPRM-Stage1Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 66.1 | — | |
| EurusPRM-Stage2Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 65.8 | — | |
| Uni-DPOModel=Qwen2.5-Math-7B, Method type=Preference Alignment, Result source=-2025.06 | 65.7 | — | |
| Uni-DPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 65.7 | — | |
| SimpleRL-Zero-7BModel=Qwen2.5-Math-7B, Method type=Reinforcement Learning, Result source=-2025.06 | 65.5 | — | |
| Dr. GRPO (Oat-Zero-7B)Model=Qwen2.5-Math-7B, Method type=Reinforcement Learning, Result source=-2025.06 | 64.9 | — | |
| GreedyDecoding Strategy=Greedy2025.05 | 64.2 | — | |
| Vanilla GRPOBase Model=Qwen2.5-7B2025.10 | 64.2 | — | |
| Scaf-GRPOBase Model=Qwen2.5-7B2025.10 | 63.8 | — | |
| Math-Shepherd-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 63.6 | — | |
| SFTModel=Qwen2.5-Math-7B, Method type=Supervised Fine-tuning, Result source=-2025.06 | 63.6 | — | |
| RAFTModel=Qwen2.5-Math-7B, Method type=Supervised Fine-tuning, Result source=-2025.06 | 63.6 | — | |
| Scaf-GRPOBase Model=Qwen2.5-Math-7B2025.10 | 63.4 | — | |
| Iterative DPOModel=Qwen2.5-Math-7B, Method type=Preference Alignment, Result source=-2025.06 | 63.4 | — | |
| Oat-ZeroBase Model=Qwen2.5-Math-7B2025.10 | 62.9 | — | |
| LUFFYBase Model=Qwen2.5-Math-7B2025.10 | 62.7 | — | |
| Vanilla GRPOBase Model=Qwen2.5-Math-7B2025.10 | 62.6 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BBase Model=DeepSeek-R1-Distill-Qwen-1.5B2025.10 | 62.1 | — | |
| SimPOModel=Qwen2.5-Math-7B, Method type=Preference Alignment, Result source=-2025.06 | 62.1 | — | |
| PPOModel=Qwen2.5-Math-7B, Method type=Reinforcement Learning, Result source=-2025.06 | 62.1 | — | |
| SimPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 62.1 | — | |
| DPOModel=Qwen2.5-Math-7B, Method type=Preference Alignment, Result source=-2025.06 | 61.6 | — | |
| DPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 61.6 | — | |
| Eurus2-PRIMEModel=Qwen2.5-Math-7B, Method type=Reinforcement Learning, Result source=-2025.06 | 61.3 | — | |
| SimpleRL-ZeroBase Model=Qwen2.5-Math-7B2025.10 | 60.8 | — | |
| CESBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 60.8 | 3,173 | |
| Uni-DPOModel=Qwen2.5-Math 1.5B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 58.2 | — | |
| Scaf-GRPOBase Model=Qwen2.5-Math-1.5B2025.10 | 57.9 | — | |
| DAPOBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 57.7 | 3,304 | |
| Vanilla GRPOBase Model=Qwen2.5-Math-1.5B2025.10 | 57.4 | — | |
| DPOModel=Qwen2.5-Math 1.5B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 56.6 | — | |
| SimPOModel=Qwen2.5-Math 1.5B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 55.3 | — | |
| BaselineModel=Qwen2.5-Math 1.5B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 53 | — | |
| Eurus2-SFTModel=Qwen2.5-Math-7B, Method type=Supervised Fine-tuning, Result source=-2025.06 | 52.2 | — | |
| Qwen2-7B-S2R-ORLBackbone=Qwen2-7B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 50.9 | — | |
| Scaf-GRPOBase Model=Llama-3.2-3B-Instruct2025.10 | 46 | — | |
| Vanilla GRPOBase Model=Llama-3.2-3B-Instruct2025.10 | 45.7 | — | |
| Llama-3.1-8B-S2R-ORLBackbone=Llama-3.1-8B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 45.2 | — | |
| BaselineModel=Qwen2.5-Math-7B, Method type=Base, Result source=-2025.06 | 44.7 | — | |
| BaselineModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 44.7 | — | |
| Qwen2.5-7BBase Model=Qwen2.5-7B2025.10 | 42.6 | — | |
| Qwen2.5-Math-7BBase Model=Qwen2.5-Math-7B2025.10 | 35.1 | — | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct2025.10 | 33.5 | — | |
| Qwen2.5-Math-1.5BBase Model=Qwen2.5-Math-1.5B2025.10 | 20 | — |