Retrieval on 375 Expert-Crafted Queries Benchmark 150 vision, 150 NLP, 75 multi-modal 1.0 (Overall)
62.6nDCG@3BLUEPRINT
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BLUEPRINTpipeline=detect→crop→OCR/IE before indexing, top_k=3, input_mode=full-page renders2026.02 | 62.6 | 60.8 | 40.7 | 43.5 | 22.2 | 71.5 | |
| LLaMA 3.2 Visionscoring=pooled embedding cosine, top_k=3, input_mode=full-page renders2026.02 | 52.1 | 49.7 | 33 | 34.2 | 15.1 | 62.3 | |
| Llama 4 Scout 17Bscoring=pooled similarity; pairwise fallback, top_k=3, input_mode=full-page renders2026.02 | 51.9 | 50.3 | 32.7 | 36.4 | 13.7 | 60.7 | |
| Pixtral 12B (2409)scoring=pooled embeddings or pairwise scoring, top_k=3, input_mode=full-page renders2026.02 | 50.2 | 48.6 | 31.4 | 35.4 | 13.1 | 59.2 | |
| PaliGemma 2scoring=pooled VL embeddings, top_k=3, input_mode=full-page renders2026.02 | 42.2 | 39.5 | 19.5 | 30.7 | 12.2 | 53.3 | |
| LLaVA 1.6 Mistral 7Bscoring=embedding head; cross-attention tie-break, top_k=3, input_mode=full-page renders2026.02 | 40 | 37.8 | 19.3 | 27.5 | 10.9 | 49.8 |