Mathematical Reasoning on AIME 2025 (Acc (%), # Tokens)
36.7Accuracy (%)ACOER
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ACOERTraining Step=1,2002026.06 | 36.7 | 8,922 | |
| GRPO-acc*Training Step=1,000, Accuracy-only GRPO=true2026.06 | 33.3 | 11,836 | |
| BaseMethod variant=Baseline2026.06 | 30 | 13,298 | |
| GRPO+LPTraining Step=400, Pre-collapse peak=true, Length Penalty (LP)=true2026.06 | 30 | 11,643 | |
| ReCutTraining Step=400, Pre-collapse peak=true2026.06 | 30 | 9,933 | |
| GRPO-LEADTraining Step=200, Pre-collapse peak=true2026.06 | 20 | 13,396 | |
| β=0 (s42)Training Step=1,200, Beta coefficient (β)=02026.06 | 10 | 8,191 | |
| GRPO-LEADTraining Step=1,2002026.06 | 10 | 4,594 | |
| ReCutTraining Step=1,2002026.06 | 10 | 4,853 | |
| GRPO+LPTraining Step=1,200, Length Penalty (LP)=true2026.06 | 0 | 1,810 |