KV Cache Efficiency on Trace-driven Simulation (No-limit Budget, 100k Pool)
21.33Hit Ratio (Low Skewness)LazyAttention
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LazyAttentionKV Cache Mem Size=No-limit, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 21.33 | 29.09 | 24.5 | |
| Block-Attention (vLLM)KV Cache Mem Size=No-limit, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 18.55 | 27.38 | 23.61 | |
| CacheBlendKV Cache Mem Size=No-limit, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 15.21 | 22.45 | 19.36 | |
| Prefix CachingKV Cache Mem Size=No-limit, Model=8B, GPU=NVIDIA H100, K=10 docs/query2026.06 | 0.58 | 2.16 | 3.25 |