Multi-modal Retrieval-Augmented Generation on MRAG-Bench
69.99AccuracyGT (Human-Annotated Image as Input)
Evaluation Results
| Method | Links | |
|---|---|---|
| GT (Human-Annotated Image as Input)Main Model=Qwen3-VL-8B, Top K=52026.05 | 69.99 | |
| GMEMain Model=Qwen3-VL-8B, Top K=1, Retriever Params=2.2B2026.05 | 64.38 | |
| SigLIP 2 So400mMain Model=MiniCPM-V4.5, Top K=4, Retriever Params=1.1B2026.05 | 63.34 | |
| UniME-V2Main Model=MiniCPM-V4.5, Top K=1, Retriever Params=7.1B2026.05 | 63.05 | |
| Ours (In-Family Surrogate)Main Model=Ovis2.5-9B, Top K=22026.05 | 61.2 | |
| Ours (Qwen3-VL-2B Surrogate)Main Model=Gemma3-12B, Top K=1, Retriever Params=2.1B2026.05 | 60.83 | |
| Zero-ShotMain Model=Qwen3-VL-8B2026.05 | 59.35 | |
| CLIP-BMain Model=Qwen3-VL-8B, Top K=1, Retriever Params=151M2026.05 | 57.8 | |
| LamRAMain Model=InternVL3.5-8B, Top K=3, Retriever Params=8B2026.05 | 47.15 |