Long-context Evaluation on RULER 32k context Average 13 tasks
0.635ScoreVanilla
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VanillaAttention Mechanism=Full Attention, Model Size=8.1B, Training Context=32k2026.05 | 0.635 | — | — | |
| Self-Pruned KVAttention Mechanism=Self-Pruned KV, tau=0.5, Model Size=8.1B, Training Context=32k2026.05 | 0.61 | -3.9 | 17.5 |