Large Language Model Serving on vLLM benchmark (128 prompts, 32 pre-fill, 256 generation tokens)
76TTFT (ms)DeInfer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeInferModel=OPT-30B, Low-rank KV cache=Enabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 76 | 39 | |
| DeInferModel=OPT-30B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 77 | 68 | |
| DeInferModel=LLaMA-65B, Low-rank KV cache=Enabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 81 | 59 | |
| DeInferModel=LLaMA-65B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 82 | 76 | |
| DeInferModel=LLaMA-3-70B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 83 | 79 | |
| DeInferModel=LLaMA-3-70B, Low-rank KV cache=Enabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 83 | 74 | |
| BaseModel=OPT-30B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 6,999 | 311 | |
| BaseModel=LLaMA-65B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 19,341 | 764 | |
| BaseModel=LLaMA-3-70B, Low-rank KV cache=Disabled, Compression ratio=40%, Hardware platform=8xA6000 (48GB), NVLink status=w/o2026.04 | 21,740 | 812 |