Mathematical Reasoning on AIME 2024 (Accuracy and Length)
56.67AccuracyARLCP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 56.67 | 6,795 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 56.45 | 5,995 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 56.25 | 8,784 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 52.08 | 6,393 | |
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 51.46 | 10,547 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 51.46 | 10,034 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 50.42 | 9,899 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 49.17 | 10,575 | |
| TACLerMode=Thinking2026.01 | 42.1 | 6,868 | |
| DeepScaleRMode=Thinking2026.01 | 40.4 | 8,565 | |
| FastCuRLMode=Thinking2026.01 | 39.8 | 10,091 | |
| TACLerMode=NoThinking2026.01 | 39.6 | 6,056 | |
| AdaptThinkMode=Efficient Reasoning2026.01 | 34.8 | 9,279 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 34.17 | 5,951 | |
| AutoThinkMode=Efficient Reasoning2026.01 | 31.7 | 8,167 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 30.83 | 6,864 | |
| STILL-3Mode=Thinking2026.01 | 30.4 | 10,605 | |
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 30 | 12,256 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 29.6 | 11,721 | |
| TLMREMode=Efficient Reasoning2026.01 | 29.2 | 8,982 | |
| O1-PrunerMode=Efficient Reasoning2026.01 | 28.9 | 10,361 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 28.75 | 7,629 | |
| OverThinkMode=Efficient Reasoning2026.01 | 28.3 | 11,269 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 27.71 | 11,937 | |
| R1-QwenMode=Thinking2026.01 | 27.7 | 12,306 | |
| R1-QwenMode=Thinking2026.01 | 27.7 | 12,306 | |
| DASTMode=Efficient Reasoning2026.01 | 26.9 | 7,745 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 26.04 | 12,181 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 25.8 | 5,386 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 23.96 | 2,671 | |
| ModelMergingMode=Efficient Reasoning2026.01 | 18.1 | 10,337 | |
| R1-QwenMode=NoThinking2026.01 | 14.8 | 4,689 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 14.38 | 4,266 |