Mathematical Reasoning on AIME 2025 (Accuracy, Length)
76.92AccuracyQwen3-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-8BPruning strategy=None2025.08 | 76.92 | 19,902.43 | |
| Qwen3-8BPruning strategy=80% lowest-entropy steps2025.08 | 76 | 11,716.63 | |
| DeepSeek-R1-14BPruning strategy=None2025.08 | 58.62 | 18,000.1 | |
| DeepSeek-R1-14BPruning strategy=80% lowest-entropy steps2025.08 | 51.72 | 10,842.07 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 39.58 | 10,692 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 39.58 | 7,560 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 39.38 | 6,410 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 38.96 | 11,384 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 38.33 | 10,793 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 37.71 | 9,622 | |
| DeepSeek-R1-7BPruning strategy=None2025.08 | 35.71 | 18,203.23 | |
| DeepSeek-R1-7BPruning strategy=80% lowest-entropy steps2025.08 | 35.71 | 11,471.17 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 35.63 | 6,584 | |
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 35 | 11,427 | |
| ARLCPBase Model=Qwen3-1.7B2026.02 | 32.92 | 9,755 | |
| DeepScaleRMode=Thinking2026.01 | 31.3 | 8,239 | |
| ARLCPBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 31.25 | 7,512 | |
| TACLerMode=Thinking2026.01 | 30.8 | 6,807 | |
| VanillaBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 29.17 | 11,189 | |
| VanillaBase Model=Qwen3-1.7B2026.02 | 28.54 | 13,287 | |
| TACLerMode=NoThinking2026.01 | 27.9 | 5,710 | |
| FastCuRLMode=Thinking2026.01 | 27.9 | 9,723 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 26.46 | 5,196 | |
| AdaptThinkMode=Efficient Reasoning2026.01 | 25.6 | 9,117 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 25.42 | 7,084 | |
| STILL-3Mode=Thinking2026.01 | 24.4 | 10,415 | |
| AutoThinkMode=Efficient Reasoning2026.01 | 23.8 | 7,647 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 23.13 | 11,879 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 22.5 | 6,626 | |
| R1-QwenMode=Thinking2026.01 | 21.5 | 12,182 | |
| R1-QwenMode=Thinking2026.01 | 21.5 | 12,182 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 21.46 | 12,258 | |
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 21.4 | 12,201 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 20.83 | 11,855 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 19.6 | 4,581 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 16.25 | 2,499 | |
| R1-QwenMode=NoThinking2026.01 | 13.3 | 4,062 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 9.79 | 3,630 |