Long-context Evaluation on RULER 8k context Average 13 tasks
79.3ScoreVanilla
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VanillaAttention Mechanism=Full Attention, Model Size=8.1B, Training Context=32k2026.05 | 79.3 | — | — | |
| Self-Pruned KVAttention Mechanism=Self-Pruned KV, tau=0.5, Model Size=8.1B, Training Context=32k2026.05 | 78.6 | -0.9 | 17.5 |