Inference Efficiency on 30k Context Length Llama-2-7B
6.6Inference Throughput (QPS)Finetuning
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| FinetuningModel=Llama-2-7B, GPU=single L40S2025.03 | 6.6 | — | — | — | |
| DBSAModel=Llama-2-7B, GPU=single L40S2025.03 | 3.4 | — | — | — | |
| Fixed ICLCaching Strategy=cached, Model=Llama-2-7B, GPU=single L40S2025.03 | 1.5 | — | — | — | |
| RetICLCaching Strategy=no cache, Model=Llama-2-7B, GPU=single L40S2025.03 | 0.8 | — | — | — | |
| DBSAModel=Llama-2-7B, Hardware=Single L40S GPU, Context Length=30k2025.03 | — | — | 3 | 0.22 | |
| FinetuningModel=Llama-2-7B, Hardware=Single L40S GPU, Fine-tuning Protocol=LoRA, Context Length=30k2025.03 | — | — | — | 0.12 | |
| Fixed ICLModel=Llama-2-7B, Hardware=Single L40S GPU, Caching Strategy=cached, Context Length=30k2025.03 | — | — | 4.5 | 0.51 | |
| RetICLModel=Llama-2-7B, Hardware=Single L40S GPU, Caching Strategy=no cache, Context Length=30k2025.03 | — | — | 1 | 1 |