Mathematical Reasoning on AIME 2025 (mean accuracy)
4.9Mean AccuracyCliff-DPO (uncertain + sampled-off)
Evaluation Results
| Method | Links | |
|---|---|---|
| Cliff-DPO (uncertain + sampled-off)evaluation_protocol=avg@64, Updated tokens=32,9162026.06 | 4.9 | |
| Qwen3-0.6B + cDPOevaluation_protocol=avg@64, Updated tokens=5,829,0522026.06 | 4.8 | |
| Cliff-DPO (all)evaluation_protocol=avg@64, Updated tokens=38,4542026.06 | 4 | |
| Cliff-DPO (uncertain)evaluation_protocol=avg@64, Updated tokens=18,1222026.06 | 3.8 | |
| Qwen3-0.6Bevaluation_protocol=avg@642026.06 | 3.5 | |
| Cliff-DPO (sampled-off)evaluation_protocol=avg@64, Updated tokens=14,7942026.06 | 3.5 | |
| Cliff-DPO (deterministic)evaluation_protocol=avg@64, Updated tokens=5,5382026.06 | 2.9 | |
| Qwen3-0.6B + DPOevaluation_protocol=avg@64, Updated tokens=2,862,8452026.06 | 2.2 |