Visual Question Answering on Enc-VQA (test)
55.1Single-Hop AccuracyLLaVA-mR2AG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-mR2AGLLM=Vicuna-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 55.1 | — | |
| Qwen2-VL-OracleLLM=Qwen2-7B, KB=Wiki, Setting=Zero-shot2025.06 | 51.2 | — | |
| ReAGGenerator=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2025.11 | 44.9 | 47 | |
| EchoSightGenerator=Mistral-7B/LLaMA-3-8B, Retriever=EVA-CLIP-8B2025.11 | 41.8 | — | |
| RMCDBase LVLM=InternVL 2.5 (38B)2026.01 | 41.4 | — | |
| ReAGGenerator=Qwen2.5-VL-3B, Retriever=EVA-CLIP-8B2025.11 | 41.3 | 42.9 | |
| Wiki-R1 7BRetrieval Model=EVA-CLIP-8B + Col., Knowledge Source=V.+ T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 41 | 37.1 | |
| ConcatBase LVLM=InternVL 2.5 (38B)2026.01 | 40.9 | — | |
| Wiki-R1 3BRetrieval Model=EVA-CLIP-8B + Col., Knowledge Source=V.+ T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 40.4 | 35.9 | |
| VLM-PRFGenerator=InternVL3-8B, Retriever=EVA-CLIP-8B2025.11 | 40.1 | 39.2 | |
| RMCDBase LVLM=BLIP-2 (T5-XXL)2026.01 | 39.5 | — | |
| mKG-RAGGenerator=LLaVA-MORE-8B, Retriever=Custom VLM2025.11 | 38.4 | 36.3 | |
| SCDBase LVLM=InternVL 2.5 (38B)2026.01 | 37.9 | — | |
| RAGBase LVLM=InternVL 2.5 (38B)2026.01 | 37.5 | — | |
| ConcatBase LVLM=BLIP-2 (T5-XXL)2026.01 | 37.4 | — | |
| VLM-PRFGenerator=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2025.11 | 37.1 | 36 | |
| ReflectiVAGenerator=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2025.11 | 36.8 | 36.8 | |
| Max ProbabilityBase LVLM=BLIP-2 (T5-XXL)2026.01 | 36.6 | — | |
| EchoSightGenerator=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B2025.11 | 36.3 | 34.2 | |
| VLM-PRFGenerator=LLaVA-MORE-8B, Retriever=EVA-CLIP-8B2025.11 | 36.3 | 35.5 | |
| RMCDBase LVLM=BLIP-2 (T5-XL)2026.01 | 36.3 | — | |
| SCDBase LVLM=BLIP-2 (T5-XXL)2026.01 | 36 | — | |
| ReflectiVAGenerator=LLaVA-MORE-8B, Retriever=EVA-CLIP-8B2025.11 | 35.5 | 35.5 | |
| ReflectiVARetrieval Model=EVA-CLIP-8B, Knowledge Source=V., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 35.5 | 35.5 | |
| Max ProbabilityBase LVLM=BLIP-2 (T5-XL)2026.01 | 35.4 | — | |
| RAGBase LVLM=BLIP-2 (T5-XXL)2026.01 | 35 | — | |
| ConcatBase LVLM=BLIP-2 (T5-XL)2026.01 | 34.9 | — | |
| SCDBase LVLM=BLIP-2 (T5-XL)2026.01 | 34.6 | — | |
| RAGBase LVLM=BLIP-2 (T5-XL)2026.01 | 34.5 | — | |
| ReflectiVAGenerator=Qwen2.5-VL-3B, Retriever=EVA-CLIP-8B2025.11 | 33.7 | 35.2 | |
| VLM-PRFGenerator=Qwen2.5-VL-3B, Retriever=EVA-CLIP-8B2025.11 | 31.1 | 32.4 | |
| DPRV+TGenerator=Multi-passage BERT, Retriever=CLIP ViT-B/322025.11 | 29.1 | — | |
| DPRV+TRetrieval Model=CLIP ViT-B/32, Knowledge Source=V. + T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 29.1 | — | |
| ReflectiVARetrieval Model=EVA-CLIP-8B, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 28 | 29.2 | |
| CoRe-MMRAGLLM=Qwen2-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 27.2 | — | |
| Max ProbabilityBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 27 | — | |
| RMCDBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 27 | — | |
| GPT-4VEvaluation Protocol=Zero-shot MLLMs2026.03 | 26.9 | 28.1 | |
| Qwen-2.5-VL 7BEvaluation Protocol=Zero-shot MLLMs2026.03 | 26.6 | 26.3 | |
| EchoSightRetrieval Model=EVA-CLIP-8B, Knowledge Source=V., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 26.4 | 24.9 | |
| SCDBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 25.7 | — | |
| ReflectiVARetrieval Model=CLIP ViT-L/14, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 24.9 | 26.7 | |
| ConsistencyBase LVLM=BLIP-2 (T5-XXL)2026.01 | 24.4 | — | |
| Qwen2-VL-1-StageLLM=Qwen2-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 24.3 | — | |
| RAGBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 23.8 | — | |
| Qwen2.5-VL-7B2025.11 | 23.6 | 23.2 | |
| Max ProbabilityBase LVLM=InternVL 2.5 (38B)2026.01 | 23.3 | — | |
| Qwen2-VL-2-StageLLM=Qwen2-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 23.1 | — | |
| ConsistencyBase LVLM=InternVL 2.5 (38B)2026.01 | 22.6 | — | |
| ConcatBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 22.6 | — | |
| EchoSightRetrieval Model=EVA-CLIP-8B, Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 22.4 | 21.7 | |
| Qwen2.5-VL-3B2025.11 | 21.9 | 21.9 | |
| ConsistencyBase LVLM=BLIP-2 (T5-XL)2026.01 | 21.4 | — | |
| Qwen2-VL-MMSTaRLLM=Qwen2-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 20.9 | — | |
| RORA-VLMLLM=Vicuna-7B, KB=Wiki+Web, Setting=Fine-tuned2025.06 | 20.3 | — | |
| CoRe-MMRAGLLM=Qwen2-7B, KB=Wiki, Setting=Zero-shot2025.06 | 20.1 | — | |
| ConsistencyBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 19.9 | — | |
| UnconditionalBase LVLM=InternVL 2.5 (38B)2026.01 | 19.8 | — | |
| EchoSightLLM=LLAMA3-8B, KB=Wiki, Setting=Fine-tuned2025.06 | 19.4 | — | |
| RMCDBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 19 | — | |
| Qwen-2.5-VL 3BEvaluation Protocol=Zero-shot MLLMs2026.03 | 18.6 | 18.8 | |
| Wiki-LLaVARetrieval Model=CLIP ViT-L/14+Con., Knowledge Source=T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | 18.3 | 19.6 | |
| Qwen2-VL-1-StageLLM=Qwen2-7B, KB=Wiki, Setting=Zero-shot2025.06 | 17.9 | — | |
| Wiki-LLaVALLM=Vicuna-7B, KB=Wiki, Setting=Fine-tuned2025.06 | 17.7 | — | |
| Wiki-LLaVAGenerator=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/14+Contriever2025.11 | 17.7 | 20.3 | |
| SCDBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 17.2 | — | |
| Qwen2-VL-2-StageLLM=Qwen2-7B, KB=Wiki, Setting=Zero-shot2025.06 | 17 | — | |
| Qwen2-VL-MMSTaRLLM=Qwen2-7B, KB=Wiki, Setting=Zero-shot2025.06 | 16.9 | — | |
| LLaVA-1.5LLM=Vicuna-7B, Setting=Zero-shot2025.06 | 16.3 | — | |
| LLaVA-v1.5-7B2025.11 | 16.3 | 16.9 | |
| LLaVA-MORE-8B2025.11 | 16 | 16.9 | |
| LLaVA-1.5 7BEvaluation Protocol=Zero-shot MLLMs2026.03 | 16 | 16.9 | |
| UnconditionalBase LVLM=BLIP-2 (T5-XXL)2026.01 | 15 | — | |
| Max ProbabilityBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 14.9 | — | |
| RAGBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 14.3 | — | |
| UnconditionalBase LVLM=LLaVA-1.5 (Vicuna 13B)2026.01 | 13.6 | — | |
| UnconditionalBase LVLM=BLIP-2 (T5-XL)2026.01 | 13.3 | — | |
| Qwen2-VL-ParamLLM=Qwen2-7B, Setting=Zero-shot2025.06 | 12.7 | — | |
| BLIP-22025.11 | 12.6 | 12.4 | |
| BLIP-2Evaluation Protocol=Zero-shot MLLMs2026.03 | 12.6 | 12.4 | |
| InstructBLIPEvaluation Protocol=Zero-shot MLLMs2026.03 | 11.9 | 12 | |
| UnconditionalBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 10.8 | — | |
| ConsistencyBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 9.6 | — | |
| ConcatBase LVLM=BLIP-2 (OPT 6.7B)2026.01 | 5.1 | — | |
| RORA-VLMGenerator=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142025.11 | — | 20.3 | |
| RORA-VLMRetrieval Model=CLIP+Google Search, Knowledge Source=V. + T., Evaluation Protocol=Retrieval-Augmented Generation2026.03 | — | 20.3 |