Long-context Retrieval on RULER 4k context (test)
95AccuracyLlama-3.1-8B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3.1-8B-Instructmax context=128k2026.05 | 95 | |
| Qwen3-8Bmax context=32k2026.05 | 94.8 | |
| gemma-3-12b-itmax context=128k2026.05 | 94.7 | |
| Ministral-3-8B-Instruct-2512-BF16max context=256k2026.05 | 94.6 | |
| Qwen3-4Bmax context=32k2026.05 | 92.7 | |
| Llama-3.2-3B-Instructmax context=128k2026.05 | 92.5 | |
| Moonlight-16B-A3B-Instructmax context=128k2026.05 | 92.2 | |
| LLaMAntino-3-ANITA-8B-Inst-DPO-ITAmax context=8k2026.05 | 91.6 | |
| gemma-3-4b-itmax context=128k2026.05 | 90.9 | |
| GPT-5 nanomax context=400k2026.05 | 87.2 | |
| Velvet-14Bmax context=128k2026.05 | 85.5 | |
| gpt-oss-20bmax context=128k2026.05 | 79.6 | |
| EngGPT2-16B-A3Bmax context=32k2026.05 | 79.5 | |
| deepseek-moe-16b-chatmax context=4k2026.05 | 77.5 | |
| FastwebMIIA-7Bmax context=16k2026.05 | 74.9 | |
| Minerva-7B-instruct-v1,0max context=4k2026.05 | 55 |