Mathematical Reasoning on Olympiad bench (Pass@1 @16, Avg. Tokens, AE Score)
57.2Pass@1 Accuracy (@16)DLER-R1-7B-Research
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLER-R1-7B-ResearchBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 57.2 | 2,316 | 1.29 | |
| Laser-DE-L4096-7BBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 56 | 2,998 | 1.09 | |
| APR-7B (Ours)Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 54 | 2,235 | 1.1 | |
| L1-Qwen-7B-MaxBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 52.5 | 2,835 | 0.89 | |
| TrainingEfficient DS-7BBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 51.8 | 4,437 | 0.55 | |
| AdaptThink-7B-delta0.05Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 50.8 | 4,402 | 0.49 | |
| DLER-R1-1.5B-ResearchBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 49.7 | 2,595 | 1.53 | |
| SB DS7B alpha 2Backbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 49.6 | 2,361 | 0.79 | |
| Laser-L8192-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 47.1 | 4,076 | 1.06 | |
| APR-1.5B (Ours)Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 46.4 | 2,450 | 1.29 | |
| L1-Qwen-1.5B-MaxBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 46.2 | 2,311 | 1.3 | |
| Original ModelBackbone=DeepSeek-R1-Distill-Qwen-7B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 46.1 | 5,395 | — | |
| Laser-DE-L4096-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 44.2 | 3,335 | 0.96 | |
| DS-1.5B-thinkprune-iter2kBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 42.9 | 3,162 | 0.88 | |
| TrainingEfficient DS-1.5BBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 40.9 | 4,523 | 0.49 | |
| AdaptThink-1.5B-delta0.05Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 40.1 | 3,612 | 0.58 | |
| Original ModelBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sampling Temperature=0.6, Max Response Length=8192, Rollout samples=162026.01 | 37.5 | 5,775 | — |