LLM Inference on LLaMA-7B v1 (serving)
12.16Decode Latency (ms/token)FlashSVD v1.5
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FlashSVD v1.5Prompt length=512, Generation length=32, Reference Baseline=HF StaticCache (SDPA)2026.05 | 12.16 | 0.43 | 2.55 | |
| FlashSVD v1.5Prompt length=512, Generation length=32, Reference Baseline=Dense KV-Cache + FA22026.05 | 12.16 | 0.43 | 2.2 | |
| FlashSVD v1.5Prompt length=2048, Generation length=128, Reference Baseline=HF StaticCache (SDPA)2026.05 | 12.24 | 1.72 | 2.5 | |
| FlashSVD v1.5Prompt length=2048, Generation length=128, Reference Baseline=Dense KV-Cache + FA22026.05 | 12.24 | 1.72 | 2.13 | |
| FlashSVD v1.5Prompt length=4096, Generation length=128, Reference Baseline=HF StaticCache (SDPA)2026.05 | 12.85 | 1.93 | 2.38 | |
| FlashSVD v1.5Prompt length=4096, Generation length=128, Reference Baseline=Dense KV-Cache + FA22026.05 | 12.85 | 1.93 | 2.07 | |
| FlashSVD v1.5Prompt length=8192, Generation length=128, Reference Baseline=HF StaticCache (SDPA)2026.05 | 14.09 | 2.39 | 2.18 | |
| FlashSVD v1.5Prompt length=8192, Generation length=128, Reference Baseline=Dense KV-Cache + FA22026.05 | 14.09 | 2.39 | 1.86 | |
| Dense KV-Cache + FA2 baselinePrompt length=2048, Generation length=1282026.05 | 26.19 | 3.51 | — | |
| Dense KV-Cache + FA2 baselinePrompt length=8192, Generation length=1282026.05 | 26.21 | 3.94 | — | |
| Dense KV-Cache + FA2 baselinePrompt length=4096, Generation length=1282026.05 | 26.52 | 3.68 | — | |
| Dense KV-Cache + FA2 baselinePrompt length=512, Generation length=322026.05 | 26.9 | 0.91 | — | |
| HF StaticCache (SDPA) baselinePrompt length=4096, Generation length=1282026.05 | 30.6 | 4.29 | — | |
| HF StaticCache (SDPA) baselinePrompt length=2048, Generation length=1282026.05 | 30.63 | 4.1 | — | |
| HF StaticCache (SDPA) baselinePrompt length=8192, Generation length=1282026.05 | 30.65 | 4.85 | — | |
| HF StaticCache (SDPA) baselinePrompt length=512, Generation length=322026.05 | 30.84 | 1.03 | — |