Long-context Evaluation on RULER 16k context Average 13 tasks
75ScoreVanilla
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VanillaAttention Mechanism=Full Attention, Model Size=8.1B, Training Context=32k2026.05 | 75 | — | — | |
| Self-Pruned KVAttention Mechanism=Self-Pruned KV, tau=0.5, Model Size=8.1B, Training Context=32k2026.05 | 74.8 | -0.3 | 17 |