Long-Context Retrieval on RULER 8K
99.67ScoreOlmo3 512swa
Evaluation Results
| Method | Links | |
|---|---|---|
| Olmo3 512swaTuning Protocol=Full param. CPT2026.07 | 99.67 | |
| HiLS-Attn NoPE-Q-CalTuning Protocol=Full param. CPT, Positional Encoding=NoPE-Q-Cal2026.07 | 99.67 | |
| HiLS-Attn HoPE-Q-CalTuning Protocol=Full param. CPT, Positional Encoding=HoPE-Q-Cal2026.07 | 99 | |
| HiLS-Attn RoPE-Q-CalTuning Protocol=Full param. CPT, Positional Encoding=RoPE-Q-Cal2026.07 | 98.67 | |
| MSA-PTTraining Strategy=from-scratch sparse pretraining2026.06 | 84.2 | |
| FullTraining Strategy=Full-Attention baseline2026.06 | 79.8 | |
| MSA-CPTTraining Strategy=sparse continued pretraining2026.06 | 77.2 | |
| Hybrid-GDNsetting=Train-from-scratch2026.04 | 76.01 | |
| Hybrid-SCAsetting=Train-from-scratch2026.04 | 75.22 | |
| Hybrid-KSAsetting=Train-from-scratch2026.04 | 73.35 | |
| Fullsetting=Train-from-scratch2026.04 | 72.85 | |
| Hybrid-SWAsetting=Train-from-scratch2026.04 | 71.69 | |
| KSAsetting=Train-from-scratch2026.04 | 65.91 | |
| LMK token tuningTuning Protocol=Freezing param.2026.07 | 22.33 | |
| Olmo3 BaseTuning Protocol=Freezing param.2026.07 | 11.34 |