Mathematical Reasoning on TabMWP (Pass@1)
91.9Pass@1Qwen2.5-7B + GRPO w/ VERL.
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 91.9 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 91.3 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 90.8 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 90.6 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 82.8 | |
| KPMath-PlusBase=DSMath, Size=7B, Zero-shot without demonstrations=true2024.03 | 78.7 | |
| KPMath-PlusBase=Qwen1.5, Size=72B, Zero-shot without demonstrations=true2024.03 | 76.7 | |
| KPMath-PlusBase=Llama-2, Size=70B, Zero-shot without demonstrations=true2024.03 | 75.1 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 72.3 | |
| KPMath-PlusBase=Llemma, Size=34B, Zero-shot without demonstrations=true2024.03 | 71.9 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 71.7 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 71.3 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 71 | |
| ChatGPTZero-shot without demonstrations=false2024.03 | 69.1 | |
| GPT-4 (0613)Zero-shot without demonstrations=false2024.03 | 67.1 | |
| KPMath-PlusBase=Mistral, Size=7B, Zero-shot without demonstrations=true2024.03 | 66.4 | |
| KPMath-PlusBase=Llama-2, Size=13B, Zero-shot without demonstrations=true2024.03 | 63.9 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 41.4 |