Mathematical Reasoning on MATH 500 (Pass@1 & Pass@8 Accuracy and Length)
62.6Accuracy Pass@1DAPO w LATR
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DAPO w LATRModel=Qwen2.5-3B, RL Algorithm=DAPO, Rollout Strategy=LATR2025.10 | 62.6 | 79 | 653 | 1,217 | |
| GRPO w LATRModel=Qwen2.5-3B, RL Algorithm=GRPO, Rollout Strategy=LATR2025.10 | 61.9 | 77.5 | 594 | 952 | |
| DAPO w Stoch.Model=Qwen2.5-3B, RL Algorithm=DAPO, Rollout Strategy=Stochastic Sampling2025.10 | 60.4 | 79.2 | 700 | 1,283 | |
| GRPO w Stoch.Model=Qwen2.5-3B, RL Algorithm=GRPO, Rollout Strategy=Stochastic Sampling2025.10 | 58.4 | 76.7 | 657 | 1,207 | |
| Qwen2.5-3BModel=Qwen2.5-3B, Configuration=Base Model2025.10 | 24.7 | 54.4 | 748 | 1,690 |