Comprehensive long-context evaluation on RULER and LongBench V2
65.29Total Average ScoreQwen2.5-14B-Instruct-1M
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-14B-Instruct-1MModel Scale=14B, Training Strategy=1M-Instruct2026.04 | 65.29 | |
| OPSDLModel Scale=32B, Training Strategy=OPSDL2026.04 | 64.93 | |
| Long-SFTModel Scale=32B, Training Strategy=Long-SFT2026.04 | 63.05 | |
| OPSDLModel Scale=14B, Training Strategy=OPSDL2026.04 | 62.3 | |
| Long-SFTModel Scale=14B, Training Strategy=Long-SFT2026.04 | 60.64 | |
| Qwen2.5-7B-Instruct-1MModel Scale=7B, Training Strategy=1M-Instruct2026.04 | 60.43 | |
| Qwen2.5-32B-InstructModel Scale=32B, Training Strategy=Base Instruct2026.04 | 59.65 | |
| Qwen2.5-14B-InstructModel Scale=14B, Training Strategy=Base Instruct2026.04 | 57.61 | |
| OPSDLModel Scale=7B, Training Strategy=OPSDL2026.04 | 56.61 | |
| LongPOModel Scale=7B, Training Strategy=LongPO2026.04 | 55.42 | |
| Long-SFTModel Scale=7B, Training Strategy=Long-SFT2026.04 | 53.99 | |
| Qwen2.5-7B-InstructModel Scale=7B, Training Strategy=Base Instruct2026.04 | 51.68 |