Mathematical Reasoning on GSM8K LONGGENBENCH
92.1Accuracy (n=15)Full Attention
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Full AttentionModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=Full Attention2025.08 | 92.1 | 91.7 | 90.8 | 91.5 | |
| RetroAttentionModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=RetroAttention2025.08 | 89.8 | 89.1 | 84.4 | 87.8 | |
| QuestModel (Backbone)=QWEN2.5-32B-INSTRUCT, Attention Mechanism=Quest2025.08 | 89.2 | 78.8 | 83.3 | 83.8 | |
| Full AttentionModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=Full Attention2025.08 | 88.7 | 84.4 | 76.6 | 83.2 | |
| RetroAttentionModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=RetroAttention2025.08 | 83.1 | 80.7 | 69.1 | 77.6 | |
| QuestModel (Backbone)=QWEN2.5-14B-INSTRUCT, Attention Mechanism=Quest2025.08 | 79 | 76.5 | 64.1 | 73.2 | |
| Full AttentionBackbone=Llama-3.1-8B-Instruct, Decoding scheme=greedy decoding2025.08 | 66.7 | 60.8 | 58 | 61.8 | |
| RetroAttentionBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 61.3 | 52.6 | 55.4 | 56.5 | |
| QuestBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 58.2 | 50.9 | 48.6 | 52.6 | |
| TOVABackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 0.2 | 0.3 | 0.2 | 0.2 | |
| StreamingLLMBackbone=Llama-3.1-8B-Instruct, KV cache budget=0.15, Retrospective window size (w)=2, Decoding scheme=greedy decoding2025.08 | 0 | 0 | 0 | 0 |