Language Understanding on MMLU LONGGENBENCH
81.8Accuracy (n=15)Full Attention
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Full AttentionModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=Full Attention2025.08 | 81.8 | 81 | 81.2 | 81.3 | |
| QuestModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=Quest2025.08 | 81 | 79.4 | 79.5 | 80 | |
| RetroAttentionModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=RetroAttention2025.08 | 80.9 | 80.3 | 79.9 | 80.4 | |
| Full AttentionModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=Full Attention2025.08 | 78.2 | 76.6 | 73.6 | 76.1 | |
| QuestModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=Quest2025.08 | 76 | 73 | 71.1 | 73.3 | |
| RetroAttentionModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=RetroAttention2025.08 | 75.8 | 73.7 | 71 | 73.5 | |
| Full AttentionBackbone=Llama-3.1-8B-Instruct, Decoding scheme=greedy decoding2025.08 | 62.5 | 58.7 | 57.1 | 59.4 | |
| RetroAttentionBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 59.3 | 55.4 | 51.2 | 55.3 | |
| QuestBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 58.8 | 54.9 | 50.6 | 54.8 | |
| TOVABackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 9.7 | 8.7 | 9.4 | 9.3 | |
| StreamingLLMBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 1.2 | 2.1 | 2.5 | 1.9 |