Visual Question Answering on InfoSeek (Unseen-Q/Unseen-E Metrics)
50.7Unseen-Q ScoreMMAgent-R2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MMAgent-R2Model=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 50.7 | 49.7 | 50.2 | |
| MMAgent-R2Model=Qwen3-VL-8B, Retriever=EVA-CLIP-8B2026.07 | 49 | 49.2 | 49.1 | |
| ReAGModel=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 48.3 | 46.2 | 47.2 | |
| Wiki-R1 7BRetrieval Model=EVA-CLIP-8B + Col., Knowledge Source=V.+ T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 47.8 | 42.3 | 44.1 | |
| Wiki-R1 3BRetrieval Model=EVA-CLIP-8B + Col., Knowledge Source=V.+ T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 46 | 40.3 | 42.2 | |
| QKVQAGenerator=LLaMA-3.1-8B, Gen. FT=true, Filter. FT=true, Support=ReflectiVA2026.01 | 45 | 43.5 | 44.2 | |
| QKVQAModel=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B2026.07 | 45 | 43.5 | 44.2 | |
| VLM-PRFModel=InternVL3-8B, Retriever=EVA-CLIP-8B2026.07 | 43.5 | 42.1 | 42.5 | |
| mKG-RAGGenerator=LLaMA-3.1-8B, Gen. FT=true, Filter. FT=true2026.01 | 41.4 | 39.6 | 40.5 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=QM-Retriever, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=true, Fine-tuning=true2025.08 | 41.4 | 39.6 | 40.5 | |
| mKG-RAGModel=LLaVA-MORE-8B, Retriever=Custom VLM2026.07 | 41.4 | 39.6 | 40.5 | |
| VLM-PRFGenerator=LLaMA-3.1-8B, Gen. FT=true, Filter. FT=true2026.01 | 41.3 | 40.6 | 40.8 | |
| OMGMGenerator=LLaMA-3.1-8B, Gen. FT=true, Filter. FT=true, Support=ReflectiVA2026.01 | 41 | 39.6 | 40.3 | |
| mR2AGLLM / MLLM=Vicuna-7B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 40.6 | 39.8 | 40.2 | |
| mR2AGModel=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142026.07 | 40.6 | 39.8 | 40.2 | |
| ReflectiVAGenerator=LLaMA-3.1-8B, Gen. FT=true, Filter. FT=true2026.01 | 40.4 | 39.8 | 40.1 | |
| ReflectiVARetrieval Model=EVA-CLIP-8B, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 40.4 | 39.8 | 40.1 | |
| ReflectiVALLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only, Fine-tuning=true2025.08 | 40.4 | 39.8 | 40.1 | |
| QKVQAGenerator=Qwen2.5-VL-7B, Gen. FT=false, Filter. FT=true2026.01 | 38.2 | 37.6 | 37.9 | |
| QKVQAModel=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 38.2 | 37.6 | 37.9 | |
| MMKB-RAGModel=Qwen2-7B, Retriever=EVA-CLIP-8B2026.07 | 36.4 | 36.3 | 36.4 | |
| OMGMGenerator=Qwen2.5-VL-7B, Gen. FT=false, Filter. FT=true, Note=Reproduction2026.01 | 35.4 | 32.9 | 34.1 | |
| QKVQAGenerator=LLaMA-3-8B, Gen. FT=false, Filter. FT=true2026.01 | 34.8 | 34.2 | 34.5 | |
| ReflectiVARetrieval Model=CLIP ViT-L/14, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 34.5 | 32.9 | 33.7 | |
| OMGMGenerator=LLaMA-3-8B, Gen. FT=false, Filter. FT=true, Note=Reproduction2026.01 | 34.2 | 32.1 | 33.1 | |
| MMhops-R1Model=Qwen2.5-VL-7B, Retriever=CLIP ViT-L/142026.07 | 33.8 | 32.6 | 33.2 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=QM-Retriever, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=true2025.08 | 32.9 | 31.3 | 32.1 | |
| CoMEMModel=Qwen2.5-VL-7B, Retriever=Custom VLM2026.07 | 32.8 | 28.5 | — | |
| Wiki-LLaVALLM / MLLM=Vicuna-7B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 30.1 | 27.8 | 28.9 | |
| Wiki-LLaVAModel=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142026.07 | 30.1 | 27.8 | 28.9 | |
| EchoSightRetrieval Model=EVA-CLIP-8B, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 30 | 30.7 | 30.4 | |
| EchoSightLLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only2025.08 | 30 | 30.7 | 30.4 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only, Fine-tuning=true2025.08 | 29.8 | 28.5 | 29.1 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only, Fine-tuning=true2025.08 | 29.4 | 27.3 | 28.3 | |
| RAG-AnythingLLM / MLLM=LLaMA-3.1-8B, Retriever=text-embedding-3-small, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false2025.08 | 28.7 | 28.1 | 28.4 | |
| Wiki-LLaVARetrieval Model=CLIP ViT-L/14+Con., Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 28.6 | 25.7 | 27.1 | |
| ReflectiVARetrieval Model=EVA-CLIP-8B, Knowledge Source=V., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 28.6 | 28.1 | 28.3 | |
| ReflectiVALLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only, Fine-tuning=true2025.08 | 28.6 | 28.1 | 28.3 | |
| RA-VQA-v2LLM / MLLM=T5-large, Retriever=ColBERT & CLIP, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 27.8 | 27.2 | 27.5 | |
| Qwen-2.5-VL 3BEvaluation Protocol=Zero-shot MLLMs2026.03 | 26.3 | 16.1 | 19.6 | |
| RA-VQALLM / MLLM=T5-large, Retriever=BERT-base, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 26.1 | 25.8 | 25.9 | |
| Qwen-2.5-VL 7BEvaluation Protocol=Zero-shot MLLMs2026.03 | 25.3 | 17.2 | 19.9 | |
| RORA-VLMRetrieval Model=CLIP+Google Search, Knowledge Source=V. + T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 25.1 | 27.3 | — | |
| RORA-VLMLLM / MLLM=Vicuna-7B, Retriever=CLIP & GS, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true2025.08 | 25.1 | 27.3 | — | |
| RORA-VLMModel=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142026.07 | 25.1 | 27.3 | — | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only2025.08 | 24.1 | 22.3 | 23.2 | |
| Qwen2.5-VL-7BGenerator=Qwen2.5-VL-7B, Gen. FT=false, Filter. FT=false2026.01 | 22.8 | 24.1 | 23.7 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only2025.08 | 21.3 | 19.8 | 20.6 | |
| Qwen3-VL-8BModel=Qwen3-VL-8B2026.07 | 20.8 | 18.5 | 19.6 | |
| Qwen2-VLLLM / MLLM=Qwen2-VL-7B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 19.8 | 18.5 | 19.2 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B2026.07 | 19.7 | 19.4 | 19.6 | |
| EchoSightRetrieval Model=EVA-CLIP-8B, Knowledge Source=V., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 18 | 19.8 | 18.8 | |
| EchoSightLLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only2025.08 | 18 | 19.8 | 18.8 | |
| GPT-4VGenerator=GPT-4V, Gen. FT=false, Filter. FT=false2026.01 | 15 | 14.3 | 14.6 | |
| GPT-4VEvaluation Protocol=Zero-shot MLLMs2026.03 | 15 | 14.3 | 14.6 | |
| GPT-4V2026.07 | 15 | 14.3 | 14.6 | |
| BLIP-2Evaluation Protocol=Zero-shot MLLMs2026.03 | 12.7 | 12.3 | 12.5 | |
| BLIP-2LLM / MLLM=Flan-T5XL, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 12.7 | 12.3 | 12.5 | |
| BLIP-2Model=Flan-T5XL2026.07 | 12.7 | 12.3 | 12.5 | |
| LLaVA-v1.5LLM / MLLM=Vicuna-7B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 9.6 | 9.4 | 9.5 | |
| LLaVA-MORELLM / MLLM=LLaMA-3.1-8B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 9 | 8.2 | 8.6 | |
| InstructBLIPEvaluation Protocol=Zero-shot MLLMs2026.03 | 8.9 | 7.4 | 8.1 | |
| InstructBLIPModel=Flan-T5XL2026.07 | 8.9 | 7.4 | 8.1 | |
| LLaVA-1.5 7BEvaluation Protocol=Zero-shot MLLMs2026.03 | 8.3 | 8.9 | 7.8 | |
| GPT-4Generator=GPT-4, Gen. FT=false, Filter. FT=false2026.01 | 7.3 | 5 | 5.9 | |
| LLaMA-3.1-8BGenerator=LLaMA-3.1-8B, Gen. FT=false, Filter. FT=false2026.01 | 2.1 | 0 | 0 | |
| LLaMA-3-8BGenerator=LLaMA-3-8B, Gen. FT=false, Filter. FT=false2026.01 | 1.5 | 0 | 0 | |
| DPRv+tModel=Multi-passage BERT, Retriever=CLIP ViT-B/322026.07 | — | — | 12.4 | |
| DPRV+TRetrieval Model=CLIP ViT-B/32, Knowledge Source=V. + T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | — | — | 12.4 | |
| EchoSightGenerator=LLaMA-3-8B, Gen. FT=false, Filter. FT=true2026.01 | — | — | 31.3 | |
| EchoSightModel=Mistral-7B/LLaMA3-8B, Retriever=EVA-CLIP-8B2026.07 | — | — | 31.3 |