LLM Generation Performance on A100 80GB (inference)
128Maximum Batch SizeLayer-Condensed KV Cache
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Layer-Condensed KV CacheModel Size=7B, Seq. Length=2048+2048, w=22024.05 | 128 | 421.02 | |
| Layer-Condensed KV CacheModel Size=7B, Seq. Length=2048+2048, w=102024.05 | 42 | 315.09 | |
| Layer-Condensed KV CacheModel Size=30B, Seq. Length=2048+2048, w=22024.05 | 32 | 108.29 | |
| LlamaModel Size=7B, Seq. Length=2048+20482024.05 | 15 | 141.1 | |
| Layer-Condensed KV CacheModel Size=30B, Seq. Length=2048+2048, w=102024.05 | 8 | 77.65 | |
| LlamaModel Size=30B, Seq. Length=2048+20482024.05 | 1 | 14.1 |