Mathematical Reasoning on Gaokao Mix 2024 (Pass@1)
35.2Pass@1Qwen2.5-7B + GRPO w/ VERL.
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 35.2 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 34.1 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 34.1 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 33 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 31.9 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 22 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 20.9 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 19.8 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 16.5 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 14.3 |