Long-context evaluation on LongBench context-dependent transfer v2
36.4Score (32k Context)Still
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| StillDecoding=deterministic greedy generation, max_new_tokens=192, repetition penalty=1.2, no-repeat 4-gram blocking=true, Judge=openai/gpt-5.52026.06 | 36.4 | 58.3 | 16.7 | |
| KV-DistillDecoding=deterministic greedy generation, max_new_tokens=192, repetition penalty=1.2, no-repeat 4-gram blocking=true, Judge=openai/gpt-5.52026.06 | 26 | 25 | 16.7 |