Mathematical Reasoning on AMC 2024 (Pass@1, Pass@32, Avg)
49.8Pass@1Qwen3-8B-Base
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-8B-BaseTraining=RLAD (ours), Context=30K2026.02 | 49.8 | 77.4 | 66.5 | |
| Qwen3-8B-BaseTraining=GRPO, Context=30K2026.02 | 49.2 | 76.9 | 61 | |
| Qwen3-8B-BaseTraining=KDRL, Context=30K2026.02 | 49.1 | 77 | 64.9 | |
| Qwen3-8B-BaseTraining=SFT, Context=30K2026.02 | 48.2 | 72.9 | 56 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 33.3 | — | — | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 33.3 | — | — | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 28.9 | — | — | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 26.7 | — | — | |
| Qwen3-1.7B-BaseTraining=RLAD (ours), Context=30K2026.02 | 24.4 | 62.4 | 39 | |
| Qwen3-1.7B-BaseTraining=KDRL, Context=30K2026.02 | 21 | 59.7 | 37.2 | |
| Qwen3-1.7B-BaseTraining=GRPO, Context=30K2026.02 | 19.7 | 58.6 | 36.5 | |
| Qwen3-8B-BaseTraining=-, Context=30K2026.02 | 18.8 | 64.6 | 36.8 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 15.6 | — | — | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 13.3 | — | — | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 11.1 | — | — | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 11.1 | — | — | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 11.1 | — | — | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 8.9 | — | — | |
| Qwen3-1.7B-BaseTraining=-, Context=30K2026.02 | 8.2 | 44.6 | 22.2 |