Visual Question Answering on E-VQA (Single-Hop/All metrics)
55.9Accuracy (Single-Hop)MMAgent-R2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MMAgent-R2Model=Qwen3-VL-8B, Retriever=EVA-CLIP-8B2026.07 | 55.9 | 54.2 | |
| MMAgent-R2Model=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 55.4 | 53.6 | |
| QKVQAModel=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 53.7 | — | |
| QKVQAModel=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B2026.07 | 49.5 | — | |
| ReAGModel=Qwen2.5-VL-7B, Retriever=EVA-CLIP-8B2026.07 | 44.9 | 47 | |
| EchoSightModel=Mistral-7B/LLaMA3-8B, Retriever=EVA-CLIP-8B2026.07 | 41.8 | — | |
| VLM-PRFModel=InternVL3-8B, Retriever=EVA-CLIP-8B2026.07 | 40.1 | 39.2 | |
| MMKB-RAGModel=Qwen2-7B, Retriever=EVA-CLIP-8B2026.07 | 39.7 | 35.9 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=QM-Retriever, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=true, Fine-tuning=true2025.08 | 38.4 | 36.3 | |
| mKG-RAGModel=LLaVA-MORE-8B, Retriever=Custom VLM2026.07 | 38.4 | 36.3 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only, Fine-tuning=true2025.08 | 36.6 | 34.9 | |
| ReflectiVALLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only, Fine-tuning=true2025.08 | 35.5 | 35.5 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only, Fine-tuning=true2025.08 | 32.9 | 31 | |
| DPRv+tModel=Multi-passage BERT, Retriever=CLIP ViT-B/322026.07 | 29.1 | — | |
| ReflectiVALLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only, Fine-tuning=true2025.08 | 28 | 29.2 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=QM-Retriever, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=true2025.08 | 27.1 | 26.1 | |
| GPT-4V2026.07 | 26.9 | 28.1 | |
| EchoSightLLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only2025.08 | 26.4 | 24.9 | |
| RAG-AnythingLLM / MLLM=LLaMA-3.1-8B, Retriever=text-embedding-3-small, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false2025.08 | 25.6 | 24.8 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true, Variant=vision-only2025.08 | 24.6 | 23.7 | |
| mKG-RAGLLM / MLLM=LLaMA-3.1-8B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only2025.08 | 24.4 | 23.4 | |
| RA-VQA-v2LLM / MLLM=T5-large, Retriever=ColBERT & CLIP, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 22.4 | 21.5 | |
| EchoSightLLM / MLLM=LLaMA-3.1-8B, Retriever=EVA-CLIP-8B, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Variant=text-only2025.08 | 22.4 | 21.7 | |
| Wiki-LLaVALLM / MLLM=Vicuna-7B, Retriever=CLIP ViT-L/14, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 21.8 | 26.4 | |
| RA-VQALLM / MLLM=T5-large, Retriever=BERT-base, Retrieval Mode (Text)=true, Retrieval Mode (Vision)=false, Fine-tuning=true2025.08 | 21.7 | 20 | |
| Qwen3-VL-8BModel=Qwen3-VL-8B2026.07 | 20.3 | 20.5 | |
| Qwen2-VLLLM / MLLM=Qwen2-VL-7B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 19.9 | 19.7 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B2026.07 | 19 | 18.8 | |
| Wiki-LLaVAModel=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142026.07 | 17.7 | 20.3 | |
| LLaVA-v1.5LLM / MLLM=Vicuna-7B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 16.3 | 16.9 | |
| LLaVA-MORELLM / MLLM=LLaMA-3.1-8B, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 15.8 | 16 | |
| BLIP-2LLM / MLLM=Flan-T5XL, Retriever=None, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=false2025.08 | 12.6 | 12.4 | |
| BLIP-2Model=Flan-T5XL2026.07 | 12.6 | 12.4 | |
| InstructBLIPModel=Flan-T5XL2026.07 | 11.9 | 12 | |
| RORA-VLMLLM / MLLM=Vicuna-7B, Retriever=CLIP & GS, Retrieval Mode (Text)=false, Retrieval Mode (Vision)=true2025.08 | — | 20.3 | |
| RORA-VLMModel=LLaVA-v1.5-7B, Retriever=CLIP ViT-L/142026.07 | — | 20.3 |