Multi-task Language Understanding on MMLU (Score, Relative Change, Gate Density)
56MMLU ScoreSelf-Pruned KV
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Self-Pruned KVAttention Mechanism=Self-Pruned KV, tau=0.5, Model Size=8.1B, Training Context=32k2026.05 | 56 | 0.1 | 28.4 | |
| VanillaAttention Mechanism=Full Attention, Model Size=8.1B, Training Context=32k2026.05 | 55.9 | — | — |