Mathematical Reasoning on AMC23 (Pass@128)
98.6Pass@128Qwen2.5-7B + PPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, Decoding Setting=Pass@1282025.09 | 98.6 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, Decoding Setting=Pass@1282025.09 | 98.4 | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true, Decoding Setting=Pass@1282025.09 | 98.3 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true, Decoding Setting=Pass@1282025.09 | 98 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, Decoding Setting=Pass@1282025.09 | 97.8 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true, Decoding Setting=Pass@1282025.09 | 95.7 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, Decoding Setting=Pass@1282025.09 | 95.4 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true, Decoding Setting=Pass@1282025.09 | 94.7 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, Decoding Setting=Pass@1282025.09 | 94.5 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, Decoding Setting=Pass@1282025.09 | 93.5 |