Mathematical Problem Solving on DAPO-Math (val)
32.5Pass@1 AccuracyDAPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DAPOsampling_strategy=LATR2025.10 | 32.5 | 54.1 | 896 | 1,880 | |
| GRPOsampling_strategy=LATR2025.10 | 28.4 | 51.9 | 853 | 1,556 | |
| DAPOsampling_strategy=Stochastic Sampling2025.10 | 26.8 | 53.1 | 1,024 | 2,022 | |
| GRPOsampling_strategy=Stochastic Sampling2025.10 | 24.1 | 51.3 | 880 | 1,732 | |
| Qwen2.5-3Bbaseline=True2025.10 | 5.6 | 20.1 | 938 | 2,203 |