Long-context reasoning on BABILong OOD
98Performance (16K Context)Trained YaRN
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Trained YaRNModel=Olmo3-7B-Instruct2026.06 | 98 | 89.5 | 78 | 67.2 | 83.2 | |
| Randomized YaRNModel=Olmo3-7B-Instruct2026.06 | 96.4 | 91.5 | 87.2 | 76.7 | 88 | |
| Randomized YaRNModel=Qwen2.5-7B-Instruct2026.06 | 93.8 | 93.4 | 90.2 | 83.9 | 90.3 | |
| LoRAModel=Qwen2.5-7B-Instruct2026.06 | 92.5 | 88.9 | 90.2 | 63 | 83.6 | |
| Trained YaRNModel=Qwen2.5-7B-Instruct2026.06 | 89.5 | 86.2 | 82.3 | 67.2 | 81.3 | |
| RPEModel=Qwen2.5-7B-Instruct2026.06 | 84.9 | 82 | 79 | 72.8 | 79.7 | |
| 0-shot + YaRNModel=Qwen2.5-7B-Instruct2026.06 | 25.6 | 18 | 20 | 21.3 | 21.2 | |
| 0-shotModel=Olmo3-7B-Instruct2026.06 | 22.6 | 19.7 | 20 | 14.1 | 19.1 | |
| 0-shot + YaRNModel=Olmo3-7B-Instruct2026.06 | 22 | 22 | 19.3 | 16.7 | 20 | |
| 0-shotModel=Qwen2.5-7B-Instruct2026.06 | 20.3 | 23.9 | 19.7 | 18.7 | 20.7 |