Long-context understanding and generation on LongBench (test)
45.1Single-Doc QA ScoreGPT-3.5-Turbo-16k
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-3.5-Turbo-16kContext window size=16k2026.03 | 45.1 | 36.2 | 23.9 | 57.6 | 51 | 54.1 | 44 | 44.5 | 44.7 | |
| Vicuna-v1.5-7B-16kContext window size=16k2026.03 | 31.8 | 18.8 | 23.2 | 56.8 | 5.3 | 47.3 | 31.9 | 26.4 | 30.5 | |
| HiCI-7B-16kContext window size=16k2026.03 | 31.1 | 26.8 | 23.6 | 57.1 | 5.8 | 62 | 36.4 | 22.7 | 33.2 | |
| HiCI-7B-16k†Context window size=16k, Inference attention mechanism=training-consistent HiCI attention during inference prefill2026.03 | 29.9 | 24.5 | 24.6 | 57 | 6.1 | 63.8 | 35.8 | 23.4 | 32.9 | |
| LongChat-7B-32kContext window size=32k2026.03 | 28.8 | 20.3 | 22.5 | 50.8 | 13 | 54.1 | 34.3 | 23.9 | 31.6 | |
| LongLoRA-7B-16kContext window size=16k2026.03 | 23.7 | 25 | 20.9 | 54.2 | 12 | 55.8 | 36.8 | 10.9 | 30.6 | |
| Llama2-7B-chat-4kContext window size=4k2026.03 | 21.7 | 18.2 | 18.5 | 49.9 | 4.1 | 48.1 | 31 | 14.3 | 26.8 |