Long-Context Reasoning on LongBench (QA, Summarization, Code Metrics)
46.2QA ScoreMLA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MLABackbone=Qwen2.5-7B, Sequence Length=128K, Evaluation Protocol=base models (not instruction-tuned), max_new_tokens=1282026.03 | 46.2 | 63.8 | 55.1 | 55 | |
| FastMKABackbone=Qwen2.5-7B, Sequence Length=128K, Evaluation Protocol=base models (not instruction-tuned), max_new_tokens=1282026.03 | 45.7 | 63.2 | 54.6 | 54.5 | |
| GQABackbone=Qwen2.5-7B, Sequence Length=128K, Evaluation Protocol=base models (not instruction-tuned), max_new_tokens=1282026.03 | 44.8 | 61.3 | 53.4 | 53.2 | |
| MHABackbone=Qwen2.5-7B, Sequence Length=128K, Evaluation Protocol=base models (not instruction-tuned), max_new_tokens=1282026.03 | 42.3 | 58.7 | 51.2 | 50.7 |