Long-context retrieval and aggregation on RULER matched-training (holdout)
48.7Performance (32k Context)Still
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| StillDecoding=deterministic greedy generation, max_new_tokens=192, repetition penalty=1.2, no-repeat 4-gram blocking=true2026.06 | 48.7 | 39.9 | 36.1 | |
| KV-DistillDecoding=deterministic greedy generation, max_new_tokens=192, repetition penalty=1.2, no-repeat 4-gram blocking=true2026.06 | 26.5 | 21.9 | 20.3 |