Long-context retrieval and reasoning on RULER 16k context
100S1 AccuracyGDN
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GDN2026.03 | 100 | 21 | 3 | 17 | 0.2 | 16 | 13 | 0.6 | 17 | 17 | 8.8 | — | — | |
| GDN-GSAKV cache ratio=50%2026.03 | 100 | 100 | 48 | 77 | 2.6 | 39 | 33 | 0 | 0.9 | 27 | 19 | — | — | |
| HAM Fixed τKV cache ratio=50%, Threshold type=Fixed2026.03 | 100 | 57 | 1.8 | 15 | 0 | 10 | 6 | 1.1 | 11 | 11 | 12 | — | — | |
| HAM Learned routerKV cache ratio=50%, Routing mechanism=Learned router2026.03 | 100 | 72 | 22 | 69 | 19 | 36 | 32 | 2.1 | 8.9 | 23 | 23 | — | — | |
| HAM Learned router + EDAKV cache ratio=50%, Routing mechanism=Learned router, Exponential decay averaging (EDA)=true2026.03 | 100 | 100 | 20 | 75 | 9.2 | 28 | 26 | 0 | 2.7 | 22 | 19 | — | — | |
| TRMContext Length=16K2026.06 | 100 | 99.8 | 98.4 | 88.2 | 71.4 | 85.7 | 85.6 | 15.6 | 52.1 | 28.8 | 33.2 | 14.4 | 64.4 | |
| YOCO (Dense)Context Length=16K, Attention=Dense2026.06 | 100 | 99.8 | 96.4 | 69.4 | 91.6 | 45.8 | 49.3 | 9.4 | 67 | 30.8 | 31.4 | 61.2 | 62.7 | |
| YOCO (CLSA)Context Length=16K, Attention=Cross-Layer Sparse Attention2026.06 | 100 | 100 | 98.4 | 70.4 | 92.4 | 53 | 47.2 | 9.8 | 61.6 | 31.2 | 32.7 | 58.4 | 62.9 | |
| HAM Learned τKV cache ratio=50%, Threshold type=Learned2026.03 | 99 | 68 | 7.8 | 31 | 0 | 16 | 21 | 0.3 | 18 | 21 | 12 | — | — | |
| Transformer2026.03 | 98 | 99 | 92 | 61 | 14 | 35 | 28 | 0.2 | 24 | 21 | 19 | — | — |