Mathematical Reasoning on AIME25 (Acc., #Tok.)
31.1AccuracyBase
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BaseBackbone=DeepScaleR-1.5B-Preview2026.02 | 31.1 | 8,301 | |
| +RLOOBackbone=DeepScaleR-1.5B-Preview2026.02 | 30.8 | 7,994 | |
| ThinkPrune-4kBackbone=DeepScaleR-1.5B-Preview2026.02 | 30.8 | 7,044 | |
| DDCABackbone=DeepScaleR-1.5B-Preview2026.02 | 29.8 | 6,109 | |
| TLMREBackbone=DeepScaleR-1.5B-Preview2026.02 | 29.2 | 6,879 | |
| +LPBackbone=DeepScaleR-1.5B-Preview2026.02 | 27.5 | 6,437 | |
| DDCABackbone=DeepSeek-R1-Distill-1.5B2026.02 | 26.5 | 9,616 | |
| ThinkPrune-4kBackbone=DeepSeek-R1-Distill-1.5B2026.02 | 25.2 | 8,879 | |
| +RLOOBackbone=DeepSeek-R1-Distill-1.5B2026.02 | 25 | 11,202 | |
| TLMREBackbone=DeepSeek-R1-Distill-1.5B2026.02 | 25 | 11,465 | |
| +LPBackbone=DeepSeek-R1-Distill-1.5B2026.02 | 23.3 | 10,160 | |
| BaseBackbone=DeepSeek-R1-Distill-1.5B2026.02 | 21.9 | 12,158 |