KV Cache Efficiency on Trace-driven Simulation (1GB Budget, 100k Pool)
3.47Low Skewness Hit RatioLazyAttention
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LazyAttentionKV Cache Mem Size=1 GB, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 3.47 | 11.11 | 13.57 | |
| Block-Attention (vLLM)KV Cache Mem Size=1 GB, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 1.84 | 6.03 | 7.27 | |
| CacheBlendKV Cache Mem Size=1 GB, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 1.51 | 4.95 | 5.96 | |
| Prefix CachingKV Cache Mem Size=1 GB, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 0 | 0 | 0 |