Long-context reasoning on MRCR OOD 8-128K
96.6Accuracy (8-16K)LoRA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LoRABase Model=Qwen2.5-7B-Instruct2026.06 | 96.6 | 74.6 | 54.5 | 31.7 | 64.4 | |
| Trained YaRNBase Model=Qwen2.5-7B-Instruct2026.06 | 69.2 | 61.9 | 61.9 | 44.1 | 59.3 | |
| Randomized YaRNBase Model=Qwen2.5-7B-Instruct2026.06 | 69.2 | 79.8 | 72.8 | 68.8 | 72.7 | |
| Trained YaRNBase Model=Olmo3-7B-Instruct2026.06 | 64.1 | 44.9 | 31.1 | 12.3 | 38.1 | |
| Randomized YaRNBase Model=Olmo3-7B-Instruct2026.06 | 62.8 | 53.1 | 39.7 | 19.5 | 43.8 | |
| RPEBase Model=Qwen2.5-7B-Instruct2026.06 | 59.2 | 67.7 | 64.6 | 52.8 | 61.1 | |
| 0-shotBase Model=Qwen2.5-7B-Instruct2026.06 | 36.5 | 46.5 | 16.5 | 5.6 | 26.3 | |
| 0-shot + YaRNBase Model=Qwen2.5-7B-Instruct2026.06 | 30.2 | 31.9 | 24.2 | 11.4 | 24.4 | |
| 0-shotBase Model=Olmo3-7B-Instruct2026.06 | 8.7 | 10.7 | 8.7 | 1 | 7.3 | |
| 0-shot + YaRNBase Model=Olmo3-7B-Instruct2026.06 | 7.5 | 12 | 9.7 | 6.5 | 8.9 |