Question Answering on SciQ (test)
92.9AccuracyFull Repetition
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full RepetitionAvg KV Cache=471.1, Avg Prefill FLOPs=3.140 T, Backbone=Qwen 2.5-3B2026.07 | 92.9 | — | |
| PARTREPAvg KV Cache=280.4, Avg Prefill FLOPs=2.481 T, Backbone=Qwen 2.5-3B, τ=0.152026.07 | 92.2 | — | |
| Naïve SummaryAvg KV Cache=277.4, Avg Prefill FLOPs=3.676 T, Backbone=Qwen 2.5-3B, Repetition Type=Appending summary2026.07 | 91.6 | — | |
| LLMLinguaAvg KV Cache=277.4, Avg Prefill FLOPs=2.105 T, Backbone=Qwen 2.5-3B, Repetition Type=Appending summary2026.07 | 90.7 | — | |
| No RepetitionAvg KV Cache=235.5, Avg Prefill FLOPs=1.549 T, Backbone=Qwen 2.5-3B2026.07 | 90.3 | — | |
| Echo EvictionAvg KV Cache=235.5, Avg Prefill FLOPs=3.152 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 90.1 | — | |
| AMOOptimizer=AMO, Model=Llama3.1-1.4B, Few-shot=02026.05 | 85.4 | — | |
| AMOOptimizer=AMO, Model=Llama3.1-760M, Few-shot=02026.05 | 81.5 | — | |
| Imagine-DeBERTa-v3-L (Retrieval)KB=Synthetic VQA+, inference_mode=Retrieval, setting=zero-shot2026.03 | 80.7 | — | |
| Imagine-DeBERTa-v3-LKB=Synthetic VQA+, setting=zero-shot2026.03 | 80.5 | — | |
| Imagine-DeBERTa-v3-LKB=Synthetic VQA, setting=zero-shot2026.03 | 78.9 | — | |
| H2O EvictionAvg KV Cache=270.8, Avg Prefill FLOPs=3.162 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 77.7 | — | |
| CAR-DeBERTa-v3-LKB=AbsAT, setting=zero-shot2026.03 | 76.9 | — | |
| CUDNNModel Scale=1B, Training Data=scientific PDFs2025.09 | 76.6 | 66.3 | |
| MuSeModel Scale=1B, Training Data=scientific PDFs2025.09 | 75.9 | 65.9 | |
| Z-LaVI (OPT-30B)setting=zero-shot2026.03 | 74 | — | |
| GPT-J-6Bsetting=zero-shot2026.03 | 73.2 | — | |
| OPT-30Bsetting=zero-shot2026.03 | 72.7 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=Commonsense-QA2024.05 | 71.8 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=RACE2024.05 | 69.2 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=BoolQ2024.05 | 67 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=ARC-Easy2024.05 | 66.6 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=Zero-shot2024.05 | 65.6 | — | |
| GPT-Neo-2.7Bsetting=zero-shot2026.03 | 64 | — | |
| Imagine-RoBERTa-LKB=Synthetic VQA, setting=zero-shot2026.03 | 63.7 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=QQP2024.05 | 60.8 | — | |
| CAR-RoBERTa-LKB=AbsAT, setting=zero-shot2026.03 | 60.7 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=SST22024.05 | 59.6 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=AG-news2024.05 | 59 | — | |
| Imagine-GPT-2-LKB=Synthetic VQA, setting=zero-shot2026.03 | 58.4 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=Conll2003-POS2024.05 | 58.2 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=MNLI2024.05 | 53.8 | — | |
| Z-LaVI (RoBERTa-L)setting=zero-shot2026.03 | 51.3 | — | |
| Z-LaVI (BART-L)setting=zero-shot2026.03 | 51 | — | |
| LLaMA-2 7B Chat (cross-task prompting)Source Task=Conll2003-NER2024.05 | 45 | — | |
| LLMLingua Comp.Avg KV Cache=78.6, Avg Prefill FLOPs=1.130 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 20.6 | — |