Mathematical Reasoning on AIME 2025 (Accuracy Avg@8)
78.1Accuracy (Avg@8)CERL full
Evaluation Results
| Method | Links | |
|---|---|---|
| CERL fullTraining=SFT+RL2026.05 | 78.1 | |
| HaloTraining=SFT+RL2026.05 | 75.1 | |
| Fixed-ratio erasure + free answerTraining=None2026.05 | 74.2 | |
| Base Qwen3-32BTraining=N/A2026.05 | 72.8 | |
| Length-penalty RLTraining=RL2026.05 | 71.1 | |
| TokenSkip†Training=SFT2026.05 | 67 |