Mathematical Reasoning on MATH-500 in-distribution (test)
79.51Average@32GRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPOModel=Qwen, Checkpoint Selection=best checkpoint2026.06 | 79.51 | 79.51 | |
| WAPOModel=Qwen, Checkpoint Selection=best checkpoint2026.06 | 78.9 | 78.9 | |
| GSPOModel=Qwen, Checkpoint Selection=best checkpoint2026.06 | 78.62 | 78.62 | |
| DAPOModel=Qwen, Checkpoint Selection=best checkpoint2026.06 | 78.01 | 78.01 | |
| GSPOModel=Smol, Checkpoint Selection=best checkpoint2026.06 | 75.88 | 75.88 | |
| DAPOModel=Smol, Checkpoint Selection=best checkpoint2026.06 | 75.66 | 75.66 | |
| GRPOModel=Smol, Checkpoint Selection=best checkpoint2026.06 | 75.01 | 75.01 | |
| WAPOModel=Smol, Checkpoint Selection=best checkpoint2026.06 | 74.86 | 74.86 | |
| GRPOModel=Gemma, Checkpoint Selection=best checkpoint2026.06 | 67.24 | 67.24 | |
| WAPOModel=Gemma, Checkpoint Selection=best checkpoint2026.06 | 66.9 | 66.9 | |
| GSPOModel=Gemma, Checkpoint Selection=best checkpoint2026.06 | 66.88 | 66.88 | |
| DAPOModel=Gemma, Checkpoint Selection=best checkpoint2026.06 | 66.5 | 66.5 |