End-to-end single-step decoding on DeepSeek-R1-Distill-LLaMA-8B 64K Context
24.1Latency (ms)LessIsMore
Evaluation Results
| Method | Links | |
|---|---|---|
| LessIsMoreToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 24.1 | |
| QuestToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 24.8 | |
| TidalDecodeToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 25.4 | |
| Full AttentionToken Budget=2K, Serving Stack=SGLang + FlashInfer, Hardware=NVIDIA A5000 GPU2025.08 | 34.4 |