End-to-end single-step decoding on DeepSeek-R1-Distill-LLaMA-8B 16K Context
23Latency (ms)LessIsMore
Evaluation Results
| Method | Links | |
|---|---|---|
| LessIsMoreToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 23 | |
| QuestToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 24.2 | |
| TidalDecodeToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 24.3 | |
| Full AttentionToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 25.3 |