Long-context retrieval on NIAH (avg)
100Score (4k Context)Qwen2.5-14B-Instruct-1M
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Qwen2.5-14B-Instruct-1MContext Window=1M2025.12 | 100 | 99.9 | 99.8 | 99.6 | 99.6 | 98.3 | |
| Llama3.1-8b-instructType=Instruct2025.12 | 99.9 | 99.6 | 99.5 | 99.6 | 99.4 | 96.9 | |
| Qwen2.5-72B2025.12 | 99.9 | 99.6 | 99.1 | 93.9 | 76.4 | 46.3 | |
| mid-2Checkpoint Stage=mid-22025.12 | 99.9 | 99.7 | 98.2 | 91 | 69 | 7.57 | |
| mid-4Checkpoint Stage=mid-42025.12 | 99.6 | 99.4 | 99.6 | 98.8 | 98.8 | 95.2 | |
| mid-3Checkpoint Stage=mid-32025.12 | 99.2 | 99.1 | 98.7 | 97.9 | 95.7 | 82.9 | |
| Qwen2.5-72BExtension=YaRN2025.12 | 97.4 | 94.3 | 92.8 | 90.4 | 85.6 | 80.1 |