Mathematical Reasoning on MAWPS (Pass@1)
97.8Pass@1Qwen2.5-7B + PPO w/ VERL.
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 97.8 | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 97.7 | |
| GPT-4 (0613)Zero-shot without demonstrations=false2024.03 | 97.6 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 97.6 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 97.3 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 96 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 96 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 95.7 | |
| KPMath-PlusBase=Qwen1.5, Size=72B, Zero-shot without demonstrations=true2024.03 | 95.5 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 95.5 | |
| KPMath-PlusBase=Llama-2, Size=70B, Zero-shot without demonstrations=true2024.03 | 95.4 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 95.4 | |
| KPMath-PlusBase=DSMath, Size=7B, Zero-shot without demonstrations=true2024.03 | 94.8 | |
| ChatGPTZero-shot without demonstrations=false2024.03 | 94.6 | |
| KPMath-PlusBase=Llemma, Size=34B, Zero-shot without demonstrations=true2024.03 | 94.5 | |
| KPMath-PlusBase=Mistral, Size=7B, Zero-shot without demonstrations=true2024.03 | 94.2 | |
| KPMath-PlusBase=Llama-2, Size=13B, Zero-shot without demonstrations=true2024.03 | 92.3 | |
| LEMMA (w/ MetaMath)Backbone=DeepSeekMath-7B, # Samples=403.59k2025.03 | 91.9 | |
| LEMMABackbone=DeepSeekMath-7B, # Samples=88.90k2025.03 | 90.9 | |
| MetaMathBackbone=DeepSeekMath-7B, # Samples=394.99k2025.03 | 90.2 | |
| GPTAugBackbone=DeepSeekMath-7B, # Samples=88.62k2025.03 | 89.6 | |
| LEMMA (w/ MetaMath)Backbone=LLaMA3-8B, # Samples=403.59k2025.03 | 89.5 | |
| RFTBackbone=DeepSeekMath-7B, # Samples=86.52k2025.03 | 89.3 | |
| ISCBackbone=DeepSeekMath-7B, # Samples=86.78k2025.03 | 89.3 | |
| MetaMathBackbone=LLaMA3-8B, # Samples=394.99k2025.03 | 88.9 | |
| LEMMABackbone=LLaMA3-8B, # Samples=88.90k2025.03 | 88.8 | |
| RefAugBackbone=LLaMA3-8B, # Samples=29.94k2025.03 | 88.4 | |
| SFTBackbone=DeepSeekMath-7B, # Samples=14.97k2025.03 | 88.1 | |
| RefAug-90kBackbone=LLaMA3-8B, # Samples=89.92k2025.03 | 87.7 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 86.9 | |
| GPTAugBackbone=LLaMA3-8B, # Samples=88.62k2025.03 | 85.9 | |
| RefAug-90kBackbone=DeepSeekMath-7B, # Samples=89.92k2025.03 | 83.1 | |
| SFTBackbone=LLaMA3-8B, # Samples=14.97k2025.03 | 83 | |
| ISCBackbone=LLaMA3-8B, # Samples=86.78k2025.03 | 82.3 | |
| S³C-Math (w/ MetaMath)Backbone=DeepSeekMath-7B, # Samples=927k2025.03 | 82.2 | |
| RefAugBackbone=DeepSeekMath-7B, # Samples=29.94k2025.03 | 82.1 | |
| RFTBackbone=LLaMA3-8B, # Samples=86.52k2025.03 | 81.8 | |
| S³C-Math (w/ MetaMath)Backbone=LLaMA3-8B, # Samples=927k2025.03 | 81.8 |