Many-shot ICL Efficiency on 30k Context Length Pool (Llama-3.1-8B)
0.08Inference Latency (Relative)Finetuning
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FinetuningModel=Llama-3.1-8B, Hardware=Single L40S GPU, Fine-tuning Protocol=LoRA, Context Length=30k2025.03 | 0.08 | — | |
| DBSAModel=Llama-3.1-8B, Hardware=Single L40S GPU, Context Length=30k2025.03 | 0.1 | 3 | |
| Fixed ICLModel=Llama-3.1-8B, Hardware=Single L40S GPU, Caching Strategy=cached, Context Length=30k2025.03 | 0.11 | 5 | |
| RetICLModel=Llama-3.1-8B, Hardware=Single L40S GPU, Caching Strategy=no cache, Context Length=30k2025.03 | 1 | 1 |