Mathematical Reasoning on Reasoning Benchmarks Overall
5.81Delta AccuracyARLCP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 5.81 | -53.05 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 4.04 | -38.69 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 2.69 | -34.96 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 2.23 | -51.47 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 2.04 | -33.97 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 1.87 | -34.68 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 1.42 | -58.1 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 1.36 | -5.21 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 0.95 | -8.64 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 0.64 | -5.69 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 0.6 | -26.32 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 0.53 | -7.01 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | -0.25 | -2.07 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | -0.69 | -3.09 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | -12.84 | -81.04 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | -17.21 | -74.59 |