Long-context Retrieval on RULER 32k context (test)
91.7AccuracyMinistral-3-8B-Instruct-2512-BF16
Evaluation Results
| Method | Links | |
|---|---|---|
| Ministral-3-8B-Instruct-2512-BF16max context=256k2026.05 | 91.7 | |
| Qwen3-8Bmax context=32k2026.05 | 91 | |
| Qwen3-4Bmax context=32k2026.05 | 89 | |
| GPT-5 nanomax context=400k2026.05 | 88.9 | |
| Llama-3.1-8B-Instructmax context=128k2026.05 | 87.3 | |
| gemma-3-12b-itmax context=128k2026.05 | 80.4 | |
| Llama-3.2-3B-Instructmax context=128k2026.05 | 77.8 | |
| gpt-oss-20bmax context=128k2026.05 | 77.1 | |
| gemma-3-4b-itmax context=128k2026.05 | 62.3 | |
| EngGPT2-16B-A3Bmax context=32k2026.05 | 42.6 | |
| FastwebMIIA-7Bmax context=16k2026.05 | 33.9 | |
| Moonlight-16B-A3B-Instructmax context=128k2026.05 | 32.7 | |
| LLaMAntino-3-ANITA-8B-Inst-DPO-ITAmax context=8k2026.05 | 28.5 | |
| deepseek-moe-16b-chatmax context=4k2026.05 | 16.9 | |
| Minerva-7B-instruct-v1,0max context=4k2026.05 | 10.1 | |
| Velvet-14Bmax context=128k2026.05 | 0 |