Long-context language understanding on LongBench (Qasper, QMSum, MultiQA, TREC, MultiNews, VCSum)
25.19Qasper ScoreBaseline
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| BaselineComp (%)=0, Backbone=LLaMA-3.1-8B-Instruct2026.06 | 25.19 | 23.2 | 92 | 39.9 | 72.5 | 26.9 | 15.91 | 42.23 | |
| STAR-KVComp (%)=50, Backbone=LLaMA-3.1-8B-Instruct2026.06 | 23.15 | 22.57 | 89.32 | 38.29 | 71.5 | 26.29 | 14.43 | 40.79 | |
| STAR-KVComp (%)=60, Backbone=LLaMA-3.1-8B-Instruct2026.06 | 22.58 | 22.4 | 81.3 | 41.04 | 65 | 25.7 | 15.6 | 39.09 | |
| PaluComp (%)=50, Backbone=LLaMA-3.1-8B-Instruct2026.06 | 15.38 | 22.1 | 73.36 | 27.58 | 63.5 | 21.66 | 1.95 | 32.21 | |
| PaluComp (%)=30, Backbone=LLaMA-3.1-8B-Instruct2026.06 | 14.47 | 23.3 | 86.71 | 26.57 | 73 | 26.19 | 8.33 | 36.93 |