Inference Efficiency on 30k Context Length (Llama-3.1-8B)
15.8Inference Throughput (QPS)Finetuning
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FinetuningModel=Llama-3.1-8B, GPU=single L40S2025.03 | 15.8 | — | |
| DBSAModel=Llama-3.1-8B, GPU=single L40S2025.03 | 12.8 | — | |
| Fixed ICLCaching Strategy=cached, Model=Llama-3.1-8B, GPU=single L40S2025.03 | 11.6 | — | |
| RetICLCaching Strategy=no cache, Model=Llama-3.1-8B, GPU=single L40S2025.03 | 1.3 | — |