Mathematical Reasoning on MATH500 (Pass@1 @16, Token Cost, AE Score)
91.8Pass@1 (Avg @16)DLER-R1-7B-Research
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLER-R1-7B-ResearchBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 91.8 | 1,429 | 0.74 | |
| Laser-DE-L4096-7BBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 91.5 | 1,634 | 0.67 | |
| APR-7B (Ours)Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 90.3 | 1,494 | 0.67 | |
| L1-Qwen-7B-MaxBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 90.2 | 2,125 | 0.47 | |
| TrainingEfficient DS-7BBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 89.1 | 2,427 | 0.34 | |
| AdaptThink-7B-delta0.05Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 88.2 | 1,946 | 0.46 | |
| DLER-R1-1.5B-ResearchBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 86.9 | 1,787 | 0.49 | |
| Original ModelBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 86.7 | 3,274 | — | |
| Original ModelBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 85.1 | 3,112 | — | |
| Laser-L8192-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 84.9 | 2,608 | 0.15 | |
| L1-Qwen-1.5B-MaxBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 84.8 | 1,899 | 0.37 | |
| APR-1.5B (Ours)Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 84.7 | 1,513 | 0.49 | |
| DS-1.5B-thinkprune-iter2kBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 83.1 | 1,822 | 0.3 | |
| Laser-DE-L4096-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 82.8 | 1,909 | 0.25 | |
| SB DS7B alpha 2Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 82.6 | 1,037 | 0.45 | |
| TrainingEfficient DS-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 81.8 | 2,285 | 0.07 | |
| AdaptThink-1.5B-delta0.05Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 80.8 | 1,532 | 0.26 |