Mathematical Reasoning on Olympiad Bench
68.13Pass@1 AccuracyMulFeRL
Evaluation Results
| Method | Links | |
|---|---|---|
| MulFeRLBackbone=Qwen3-4B-Inst, Training Strategy=Multi-turn Feedback-guided Reinforcement Learning2026.01 | 68.13 | |
| Critique-GRPO (CoT Critique)Backbone=Qwen3-8B2025.06 | 66.8 | |
| Critique-GRPO (Self-Critique)w/ External Supervision=true2025.06 | 66.2 | |
| R1-GRPOw/ External Supervision=true2025.06 | 65.6 | |
| Critique-GRPO (Self-Critique & Self-Evaluation)w/ External Supervision=false2025.06 | 65.5 | |
| GPT-o1-miniDecoding=Greedy2025.02 | 65.3 | |
| Critique-GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 62.4 | |
| Dr.GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 60.36 | |
| Pass@8 (Upper Bound)Decoding Strategy=Pass@82025.05 | 59.4 | |
| GRPOBackbone=Qwen3-4B-Inst, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 59.05 | |
| SFTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 49.17 | |
| Critique-GRPOTraining Data Volume=4k, Critique Mode=CoT-Critique2025.06 | 48.6 | |
| CITL-FTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 48.55 | |
| RAFTBackbone=Qwen3-4B-Inst, Training Strategy=Supervised Learning-based Finetuning2026.01 | 48.34 | |
| SCOPEDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 46.8 | |
| SFTw/ External Supervision=true2025.06 | 46.4 | |
| Majority @8Decoding Strategy=Majority @82025.05 | 46.2 | |
| SCOPE (Step Skipping)Decoding Strategy=Best-of-8 strategy, Ablation=Step Skipping2025.05 | 45.8 | |
| SCOPE (w/o AST normalization)Decoding Strategy=Best-of-8 strategy, Ablation=w/o AST normalization2025.05 | 45.6 | |
| SCOPE (w/o code translation)Decoding Strategy=Best-of-8 strategy, Ablation=w/o code translation2025.05 | 45.5 | |
| SCOPE (Step Replacement)Decoding Strategy=Best-of-8 strategy, Ablation=Step Replacement2025.05 | 45.3 | |
| Qwen3-4B-InstBackbone=Qwen3-4B-Inst, Training Strategy=Base Model2026.01 | 45.13 | |
| Qwen2.5-Math-7B-S2R-ORLBackbone=Qwen2.5-Math-7B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 44.9 | |
| RLHFlow-PRM-Deepseek-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 44.3 | |
| Qwen3-8B (w/ Think)2025.06 | 44.1 | |
| Qwen3-8BBackbone=Qwen3-8B2025.06 | 44.1 | |
| Qwen2.5-Math-7B-PRM800KDecoding Strategy=Best-of-8 strategy, Training Data=manually annotated (PRM800K)2025.05 | 43.5 | |
| Oat-ZeroTraining Data Volume=46k2025.06 | 43.4 | |
| GPT-4oDecoding=Greedy2025.02 | 43.3 | |
| RLHFlow-PRM-Mistral-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 43.3 | |
| GPT-4o-2024-08-06Synthesis Model=-2024.10 | 43.3 | |
| Skywork-PRM-1.5BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 43 | |
| Skywork-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 42.8 | |
| MulFeRLBackbone=Qwen2.5-7B-Base, Training Strategy=Multi-turn Feedback-guided Reinforcement Learning2026.01 | 42.49 | |
| Critique-GRPO (CoT Critique)Backbone=Qwen2.5-7B-Base2025.06 | 42.4 | |
| VI-CuRLBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Reward Supervision=Verifier-Free, Verifier Strategy=Majority Vote2026.02 | 42 | |
| PRIME-ZeroTraining Data Volume=46k2025.06 | 40.3 | |
| EurusPRM-Stage1Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 40.2 | |
| Progressive Thought EncodingBackbone Model=DeepSeek-R1-Distill-Llama-8B, Maximum TFLOPs of Attention=4.6, Peak GPU Mem. (%)=59.8, Mean GPU Mem. (%)=46.82026.02 | 39.7 | |
| Math-Shepherd-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 39.1 | |
| LoRABackbone Model=Qwen2.5-7B-Instruct, Maximum TFLOPs of Attention=5.7, Peak GPU Mem. (%)=85.8, Mean GPU Mem. (%)=59.32026.02 | 38.7 | |
| Progressive Thought EncodingBackbone Model=Qwen2.5-7B-Instruct, Maximum TFLOPs of Attention=3.6, Peak GPU Mem. (%)=67.2, Mean GPU Mem. (%)=48.62026.02 | 38.7 | |
| Critique-GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 38.64 | |
| EurusPRM-Stage2Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 38.5 | |
| Qwen2-Math-7B-ScaleQuestBase Model=Qwen2-Math-7B, Synthesis Model=Qwen2-Math-7B-Instruct2024.10 | 38.5 | |
| GreedyDecoding Strategy=Greedy2025.05 | 38.2 | |
| Qwen2-Math-7B-InstructSynthesis Model=-2024.10 | 37.8 | |
| GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 36.77 | |
| MathScaleBackbone=Qwen2.5-7B-Instruct, Problem Crafter=ChatGPT, Chain-of-Thought=true, Zero-shot=true2025.04 | 36.4 | |
| RV-SynBackbone=Qwen2.5-7B-Instruct, Problem Crafter=7B / 72B model, Chain-of-Thought=true, Zero-shot=true2025.04 | 36.4 | |
| Official ModelBackbone=Qwen2.5-7B-Instruct, Problem Crafter=-, Chain-of-Thought=true, Zero-shot=true2025.04 | 36 | |
| LoRAcBackbone Model=Qwen2.5-7B-Instruct, Maximum TFLOPs of Attention=3.5, Peak GPU Mem. (%)=63.1, Mean GPU Mem. (%)=45.42026.02 | 35.9 | |
| Dr.GRPOBackbone=Qwen2.5-7B-Base, Training Strategy=Reinforcement Learning-based Finetuning2026.01 | 35.79 | |
| Orca-MathBackbone=Qwen2.5-7B-Instruct, Problem Crafter=GPT-4, Chain-of-Thought=true, Zero-shot=true2025.04 | 35.4 | |
| LoRABackbone Model=DeepSeek-R1-Distill-Llama-8B, Maximum TFLOPs of Attention=7.4, Peak GPU Mem. (%)=88.7, Mean GPU Mem. (%)=53.52026.02 | 35.3 | |
| MetaMathBackbone=Qwen2.5-7B-Instruct, Problem Crafter=ChatGPT, Chain-of-Thought=true, Zero-shot=true2025.04 | 35 | |
| VI-CuRLBackbone=Qwen2.5-Math-1.5B, Reward Supervision=Oracle Reward, Verifier Strategy=w. Verifier2026.02 | 34.8 | |
| PromptCoTBackbone=Qwen2.5-7B-Instruct, Problem Crafter=72B Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 34.8 | |
| Mammoth2Backbone=Qwen2.5-7B-Instruct, Problem Crafter=72B Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 34.7 | |
| Jiuzhang3.0Backbone=Qwen2.5-7B-Instruct, Problem Crafter=7B Math Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 34.7 | |
| BaselineBackbone Model=Qwen2.5-7B-Instruct2026.02 | 34.7 | |
| SimpleRL-ZeroTraining Data Volume=46k2025.06 | 34.7 | |
| ScaleQuestBackbone=Qwen2.5-7B-Instruct, Problem Crafter=7B Math Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 34.4 | |
| Qwen2-Math-7B-Numina-MathBase Model=Qwen2-Math-7B, Synthesis Model=GPT-4o2024.10 | 33.6 | |
| LoRAcBackbone Model=DeepSeek-R1-Distill-Llama-8B, Maximum TFLOPs of Attention=4.6, Peak GPU Mem. (%)=59.1, Mean GPU Mem. (%)=47.12026.02 | 31.9 | |
| CITL-FTBackbone=Qwen2.5-7B-Base, Training Strategy=Supervised Learning-based Finetuning2026.01 | 30.89 | |
| Qwen2.5-7B-BaseBackbone=Qwen2.5-7B-Base2025.06 | 30.4 | |
| DeepSeekMath-7B-ScaleQuestBase Model=DeepSeekMath-7B, Synthesis Model=Qwen2-Math-7B-Instruct2024.10 | 29.9 | |
| RAFTBackbone=Qwen2.5-7B-Base, Training Strategy=Supervised Learning-based Finetuning2026.01 | 29.58 | |
| Progressive Thought EncodingBackbone Model=Qwen2.5-3B-Instruct, Maximum TFLOPs of Attention=2.7, Peak GPU Mem. (%)=45.3, Mean GPU Mem. (%)=32.62026.02 | 29 | |
| BaselineBackbone Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 28.7 | |
| Qwen2.5-7B-BaseBackbone=Qwen2.5-7B-Base, Training Strategy=Base Model2026.01 | 28.13 | |
| LoRABackbone Model=Qwen2.5-3B-Instruct, Maximum TFLOPs of Attention=4.2, Peak GPU Mem. (%)=82.8, Mean GPU Mem. (%)=63.52026.02 | 27.8 | |
| LoRAcBackbone Model=Qwen2.5-3B-Instruct, Maximum TFLOPs of Attention=2.6, Peak GPU Mem. (%)=38, Mean GPU Mem. (%)=312026.02 | 27.7 | |
| SFTBackbone=Qwen2.5-7B-Base, Training Strategy=Supervised Learning-based Finetuning2026.01 | 27.27 | |
| BaselineBackbone Model=Qwen2.5-3B-Instruct2026.02 | 27.2 | |
| Mistral-7B-ScaleQuestBase Model=Mistral-7B, Synthesis Model=Qwen2-Math-7B-Instruct2024.10 | 26.8 | |
| Qwen2-7B-S2R-ORLBackbone=Qwen2-7B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 26.2 | |
| Llama3-8B-ScaleQuestBase Model=Llama3-8B, Synthesis Model=Qwen2-Math-7B-Instruct2024.10 | 25.3 | |
| Qwen2-Math-7B-DART-MathBase Model=Qwen2-Math-7B, Synthesis Model=DSMath-7B-RL2024.10 | 23.1 | |
| RV-SynBackbone=Phi-3-mini, Problem Crafter=7B / 72B model, Chain-of-Thought=true, Zero-shot=true2025.04 | 22.4 | |
| ScaleQuestBackbone=Phi-3-mini, Problem Crafter=7B Math Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 21.8 | |
| DeepSeekMath-7B-DART-MathBase Model=DeepSeekMath-7B, Synthesis Model=DSMath-7B-RL2024.10 | 21.7 | |
| Mammoth2Backbone=Phi-3-mini, Problem Crafter=72B Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 21.3 | |
| PromptCoTBackbone=Phi-3-mini, Problem Crafter=72B Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 21.3 | |
| Llama-3.1-8B-S2R-ORLBackbone=Llama-3.1-8B, Training Strategy=Outcome-level RL, Decoding=Greedy2025.02 | 20.7 | |
| MetaMathBackbone=Phi-3-mini, Problem Crafter=ChatGPT, Chain-of-Thought=true, Zero-shot=true2025.04 | 20.7 | |
| MathScaleBackbone=Phi-3-mini, Problem Crafter=ChatGPT, Chain-of-Thought=true, Zero-shot=true2025.04 | 20.3 | |
| DeepSeekMath-7B-Numina-MathBase Model=DeepSeekMath-7B, Synthesis Model=GPT-4o2024.10 | 19.9 | |
| Orca-MathBackbone=Phi-3-mini, Problem Crafter=GPT-4, Chain-of-Thought=true, Zero-shot=true2025.04 | 19.6 | |
| Jiuzhang3.0Backbone=Phi-3-mini, Problem Crafter=7B Math Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 19.6 | |
| Mistral-7B-NuminaMathBase Model=Mistral-7B, Synthesis Model=GPT-4o2024.10 | 19.4 | |
| DeepSeekMath-7B-RLSynthesis Model=-2024.10 | 19 | |
| Qwen2-Math-7B-MetaMathBase Model=Qwen2-Math-7B, Synthesis Model=GPT-3.52024.10 | 17.9 | |
| Llama3-8B-NuminaMathBase Model=Llama3-8B, Synthesis Model=GPT-4o2024.10 | 17.8 | |
| Qwen2.5-Math-7B-BaseTraining Data Volume=None2025.06 | 17.6 | |
| RV-SynBackbone=LLaMA-3-8B-Instruct, Problem Crafter=7B / 72B model, Chain-of-Thought=true, Zero-shot=true2025.04 | 16.4 | |
| Jiuzhang3.0Backbone=LLaMA-3-8B-Instruct, Problem Crafter=7B Math Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 15.9 | |
| PromptCoTBackbone=LLaMA-3-8B-Instruct, Problem Crafter=72B Model, Chain-of-Thought=true, Zero-shot=true2025.04 | 15.7 | |
| Official ModelBackbone=Phi-3-mini, Problem Crafter=-, Chain-of-Thought=true, Zero-shot=true2025.04 | 15.7 |