Mathematical Reasoning on AIME 2025 (pass@1 Last, pass@1 Best)
32.7Pass@1 (Last Sample)SDPG-UFKL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SDPG-UFKLBackbone=Qwen3-4B, Sampling Strategy=mean@32, Training Steps=4002026.06 | 32.7 | 33.5 | |
| SDPG-URKLBackbone=Qwen3-4B, Sampling Strategy=mean@32, Training Steps=4002026.06 | 30.7 | 30.8 | |
| RLSDBackbone=Qwen3-4B, Sampling Strategy=mean@32, Training Steps=4002026.06 | 30 | 30.4 | |
| GRPOBackbone=Qwen3-4B, Sampling Strategy=mean@32, Training Steps=4002026.06 | 24.2 | 27.9 |