Time to First Token (TTFT) measurement on Context sequences 32K-256K (vLLM 0.18.0/0.20.0)
2.944TTFT (Dense, s)Qwen3.5-9B (distilled indexer)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3.5-9B (distilled indexer)Indexer=distilled, Hardware=GPU 80G, Serving=vLLM 0.20 chunk 4K, Dense baseline=FA2, Context=32K, top-k=10242026.06 | 2.944 | 2.747 | 1.07 | |
| Qwen3.5-0.8B (distilled indexer)Indexer=distilled, Hardware=NPU 64G, Serving=vLLM 0.18 full, Dense baseline=Ascend FA, Context=128K, top-k=10242026.06 | 5.957 | 4.613 | 1.29 | |
| Qwen3.5-9B (distilled indexer)Indexer=distilled, Hardware=GPU 80G, Serving=vLLM 0.20 chunk 4K, Dense baseline=FA2, Context=64K, top-k=20482026.06 | 6.863 | 6.024 | 1.14 | |
| Qwen3.5-9B (distilled indexer)Indexer=distilled, Hardware=GPU 80G, Serving=vLLM 0.20 chunk 4K, Dense baseline=FA2, Context=100K, top-k=20482026.06 | 11.78 | 9.173 | 1.28 | |
| Qwen3.5-9B (distilled indexer)Indexer=distilled, Hardware=GPU 80G, Serving=vLLM 0.20 chunk 4K, Dense baseline=FA2, Context=128K, top-k=20482026.06 | 17.652 | 12.452 | 1.42 | |
| Qwen3.5-0.8B (distilled indexer)Indexer=distilled, Hardware=NPU 64G, Serving=vLLM 0.18 full, Dense baseline=Ascend FA, Context=256K, top-k=10242026.06 | 17.928 | 10.491 | 1.71 | |
| Qwen3.5-9B (distilled indexer)Indexer=distilled, Hardware=GPU 80G, Serving=vLLM 0.20 chunk 4K, Dense baseline=FA2, Context=256K, top-k=20482026.06 | 50.915 | 26.322 | 1.93 |