Mathematical Reasoning on GSM8K (Acc, Time, Peak, Dep)
88.32Accuracy (GSM8K)LThinker++
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LThinker++Model Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 88.32 | 12.7 | 408 | 300,000 | |
| TokenSkipModel Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 87.92 | 47.8 | 775 | 1,000,000 | |
| VanillaModel Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 87.34 | 12.7 | 812 | 1,100,000 | |
| CoTModel Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 86.12 | 99.6 | 513 | 100,000 | |
| CoTModel Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 85.14 | 129 | 550 | 200,000 | |
| LThinker*Model Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 84.94 | 13.5 | 376 | 300,000 | |
| VanillaModel Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 82.79 | 16.1 | 811 | 1,300,000 | |
| LThinker++Model Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 82.23 | 13.3 | 424 | 300,000 | |
| Distill-R1Model Series=Qwen2.5-7B, Serving Setting=Throughput2026.04 | 81.88 | 336 | 844 | 1,100,000 | |
| TokenSkipModel Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 79.4 | 54.1 | 838 | 1,200,000 | |
| LThinker*Model Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 75.54 | 12.5 | 357 | 200,000 | |
| Distill-R1Model Series=Llama3.1-8B, Serving Setting=Throughput2026.04 | 73.62 | 154.8 | 395 | 100,000 |