Visual Question Answering on OK-VQA (test)
76.5AccuracyMATA
Evaluation Results
| Method | Links | |
|---|---|---|
| MATAType=Compositional, Agentic types=multi-agent, Mode=Domain-Specific2026.01 | 76.5 | |
| MATAType=Compositional, Agentic types=multi-agent, Mode=General2026.01 | 76 | |
| InternVL3.5Type=Monolithic, Agentic types=non-agentic/non-specified, Model Size=8B2026.01 | 75.7 | |
| InternVL2.5Type=Monolithic, Agentic types=non-agentic/non-specified, Model Size=8B2026.01 | 75.2 | |
| InternVL3Type=Monolithic, Agentic types=non-agentic/non-specified, Model Size=8B2026.01 | 74.7 | |
| Qwen2.5-VLType=Monolithic, Agentic types=non-agentic/non-specified, Model Size=7B2026.01 | 71.8 | |
| MAD-RAGBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 71.32 | |
| MAD-RAGBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 70.22 | |
| CADBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 69.38 | |
| CADBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 68.44 | |
| ALFARBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 67.66 | |
| OPERABackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 67.55 | |
| AdaCADBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 67.01 | |
| Vanilla RAGBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 66.76 | |
| OPERABackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 66.47 | |
| MAD-RAGBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 66.44 | |
| CADBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 66.24 | |
| AdaCADBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 66.15 | |
| DoLABackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 66.02 | |
| Prophet++ (mPLUG)Knowledge Resource=GPT-4o API, Base VQA Model=mPLUG2023.03 | 65.7 | |
| DoLABackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 65.65 | |
| SPINBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 65.58 | |
| VCDBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 65.55 | |
| Vanilla RAGBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 65.46 | |
| CADBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 65.45 | |
| ALFARBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 65.32 | |
| MAD-RAGBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 64.91 | |
| PaLI#Params=17B, #Pre. Data=1.6B2023.05 | 64.5 | |
| PALIProtocol=Supervised, Configuration=OK-VQA, finetune2023.06 | 64.5 | |
| PALIKnowledge Resource=multimodal pretraining, Model Parameters=17B2023.03 | 64.5 | |
| VCDBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 64.48 | |
| Closed-book (parametric)Backbone=LLaVA-1.5-13B, Context Setting=Parametric only2026.01 | 64.32 | |
| VHRBackbone=LLaVA-1.5-13B, Context Setting=Oracle Contexts2026.01 | 64.13 | |
| Vanilla RAGBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 63.8 | |
| DoLABackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 63.77 | |
| ALFARBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 63.55 | |
| VCDBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 63.17 | |
| DWIMType=Compositional, Agentic types=single-agent2026.01 | 62.8 | |
| ALFARBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 62.63 | |
| Prophet (mPLUG)Knowledge Resource=GPT-3 API, Base VQA Model=mPLUG2023.03 | 62.5 | |
| MMA-RAGModel=Qwen2VL-7B2026.02 | 62.4 | |
| Vanilla RAGBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 62.24 | |
| RIRModel=Qwen2VL-7B2026.02 | 62.2 | |
| DoLABackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 62.09 | |
| VHRBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 61.99 | |
| Closed-book (parametric)Backbone=Qwen2.5-VL-7B, Context Setting=Parametric only2026.01 | 61.92 | |
| Closed-book (parametric)Backbone=LLaVA-1.5-7B, Context Setting=Parametric only2026.01 | 61.65 | |
| SPINBackbone=LLaVA-1.5-7B, Context Setting=Oracle Contexts2026.01 | 61.55 | |
| Prophet (MCAN)Knowledge Resource=GPT-3 API, Base VQA Model=MCAN2023.03 | 61.1 | |
| VCDBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 60.81 | |
| PromptCapKnowledge Resource=GPT-3 API, Base VQA Model=OFA, GPT-3 training query=true2023.03 | 60.4 | |
| AVISProtocol=Few-shot, Decision logic=Dynamic decision making2023.06 | 60.2 | |
| MMA-RAGModel=Idefics2-8B2026.02 | 60.1 | |
| Raw dataMLLM=LLaVA2026.03 | 60.1 | |
| Closed-book (parametric)Backbone=Qwen2.5-VL-3B, Context Setting=Parametric only2026.01 | 59.97 | |
| HYDRAType=Compositional, Agentic types=non-agentic/non-specified2026.01 | 59.4 | |
| AdaCADBackbone=Qwen2.5-VL-3B, Context Setting=Oracle Contexts2026.01 | 59.11 | |
| REVEALKnowledge Sources=WIT + CC12M + Wikidata + VQA-2, Memory (GB)=4.2 + 10 + 9932022.12 | 59.1 | |
| REVEALProtocol=Supervised2023.06 | 59.1 | |
| AdaCADBackbone=Qwen2.5-VL-7B, Context Setting=Oracle Contexts2026.01 | 58.84 | |
| PromptCapEvaluation Setting=Supervised, Parameters=175B, Use Extra PLM?=false, With extra V-L Pre-training?=true2023.05 | 58.8 | |
| TwOKnowledge in Input Text=Wikipedia+Frozen OFA (0.93B)+Frozen GPT-3 (175B), Ensemble=true2023.05 | 58.72 | |
| Zero shotModel=Qwen2VL-7B2026.02 | 58.7 | |
| Few shotModel=Idefics2-8B2026.02 | 58.7 | |
| Few shotModel=Qwen2VL-7B2026.02 | 58.5 | |
| AVISProtocol=Few-shot, Ablation=without Object2023.06 | 58.3 | |
| MMA-RAGModel=Idefics3-8B2026.02 | 58.3 | |
| ReVIVEModel Configuration=Ensemble, Knowledge Sources=Wikidata + Frozen GPT-3, Memory (GB)=4.6 + 354 + 5002022.12 | 58 | |
| REVEAL-LargeKnowledge Sources=WIT + CC12M + Wikidata + VQA-2, Memory (GB)=2.8 + 10 + 9932022.12 | 58 | |
| REVIVEKnowledge in Input Text=Wikidata+Frozen GPT-3 (175B), Ensemble=true2023.05 | 58 | |
| ReVIVEProtocol=Supervised2023.06 | 58 | |
| Flamingo#Params=80B, #Pre. Data=2.3B2023.05 | 57.8 | |
| FlamingoProtocol=Few-shot2023.06 | 57.8 | |
| FlamingoKnowledge Resource=multimodal pretraining, Model Parameters=80B2023.03 | 57.8 | |
| OursMLLM=LLaVA2026.03 | 57.6 | |
| TwOKnowledge in Input Text=Wikipedia+Frozen OFA (0.93B)+Frozen GPT-3 (175B), Ensemble=false2023.05 | 57.57 | |
| RIRModel=Idefics2-8B2026.02 | 56.7 | |
| MAILModel Inputs=Question + Image, External Knowledge=Frozen MiniGPT-4 (7B)* + ConceptNet, Fusion Strategy=Modality-aware2024.02 | 56.69 | |
| ours-ViB#Params=0.88B (+0.93B), #Pre. Data=0.44M, Base Model=ViB2023.05 | 56.67 | |
| ReVIVEModel Configuration=Single, Knowledge Sources=Wikidata + Frozen GPT-3, Memory (GB)=1.5 + 354 + 5002022.12 | 56.6 | |
| REVIVEKnowledge in Input Text=Wikidata+Frozen GPT-3 (175B), Ensemble=false2023.05 | 56.6 | |
| REVIVEKnowledge Resource=GPT-3 API, GPT-3 training query=true2023.03 | 56.6 | |
| RIRModel=Idefics3-8B2026.02 | 56.6 | |
| PaLI#Params=15B, #Pre. Data=1.6B2023.05 | 56.5 | |
| TwOKnowledge in Input Text=Wikipedia+Frozen OFA (0.93B), Ensemble=true2023.05 | 56.49 | |
| ours-LXM#Params=0.98B (+0.93B), #Pre. Data=0.44M, Base Model=LXMERT2023.05 | 56.49 | |
| PerceptionGPT2023.11 | 56.2 | |
| CLIPModel=Qwen2VL-7B2026.02 | 56 | |
| Obf-WeakBlurMLLM=LLaVA2026.03 | 55.9 | |
| DownsampleMLLM=LLaVA2026.03 | 55.7 | |
| TwOKnowledge in Input Text=Wikipedia+Frozen OFA (0.93B), Ensemble=false2023.05 | 55.33 | |
| REVEAL-BaseKnowledge Sources=WIT + CC12M + Wikidata + VQA-2, Memory (GB)=0.8 + 7.5 + 7442022.12 | 55.2 | |
| AVISProtocol=Few-shot, Ablation=without Search2023.06 | 55 | |
| CLIPModel=Idefics2-8B2026.02 | 54.8 | |
| KATKnowledge Sources=Wikidata + GPT-3, Approx. Params=175B2022.10 | 54.41 | |
| KAT (Ensemble)Model Inputs=Question + Caption + Object Tags, External Knowledge=Frozen GPT-3 (175B) + Wikidata, Fusion Strategy=Modality-agnostic2024.02 | 54.41 | |
| KATModel Configuration=Ensemble, Knowledge Sources=Wikidata + Frozen GPT-3, Memory (GB)=4.6 + 352 + 5002022.12 | 54.4 | |
| KATKnowledge in Input Text=Wikidata+Frozen GPT-3 (175B), Ensemble=true2023.05 | 54.4 | |
| KATProtocol=Supervised2023.06 | 54.4 | |
| Obf-StrongBlurMLLM=LLaVA2026.03 | 54.4 |