External Knowledge-dependent Image Question Answering on OK-VQA
91.9AccuracyStaR-KVQA_Gemma
Evaluation Results
| Method | Links | |
|---|---|---|
| StaR-KVQA_GemmaModel Inputs=Question + Image, External Knowledge=Fine-tuned Gemma-3-12B2025.10 | 91.9 | |
| StaR-KVQA_QwenModel Inputs=Question + Image, External Knowledge=Fine-tuned Qwen2.5-VL-7B2025.10 | 91.51 | |
| StaR-KVQA_LlamaModel Inputs=Question + Image, External Knowledge=Fine-tuned Llama-3.2-11B-Vision2025.10 | 90.01 | |
| SDFTModel Inputs=Question + Image, External Knowledge=Fine-tuned Qwen2.5-VL-7B2025.10 | 82.56 | |
| Qwen2.5-VL-72BModel Inputs=Question + Image, External Knowledge=Qwen2.5-VL-72B2025.10 | 80.75 | |
| Gemini 2.5 ProModel Inputs=Question + Image, External Knowledge=Gemini 2.5 Pro2025.10 | 80.53 | |
| Gemini 2.5 FlashModel Inputs=Question + Image, External Knowledge=Gemini 2.5 Flash2025.10 | 79.97 | |
| CoT + SFTModel Inputs=Question + Image, External Knowledge=Fine-tuned Qwen2.5-VL-7B2025.10 | 79.58 | |
| Gemma-3-27BModel Inputs=Question + Image, External Knowledge=Gemma-3-27B2025.10 | 79.34 | |
| M2-ReasoningModel Inputs=Question + Image, External Knowledge=M2-Reasoning-7B2025.10 | 78.63 | |
| GPT-4oModel Inputs=Question + Image, External Knowledge=GPT-4o2025.10 | 77.86 | |
| CoTModel Inputs=Question + Image, External Knowledge=Qwen2.5-VL-7B2025.10 | 76.88 | |
| LLaVA-CoTModel Inputs=Question + Image, External Knowledge=Fine-tuned Llama-3.2-11B-Vision2025.10 | 76.57 | |
| SFTModel Inputs=Question + Image, External Knowledge=Fine-tuned Qwen2.5-VL-7B2025.10 | 76.36 | |
| Qwen2.5-VL-7BModel Inputs=Question + Image, External Knowledge=Qwen2.5-VL-7B2025.10 | 75.74 | |
| Gemma-3-12BModel Inputs=Question + Image, External Knowledge=Gemma-3-12B2025.10 | 71.4 | |
| PaliGemma2-28BResolution=448x448, Fine-tuning=per-task2025.12 | 70.6 | |
| PaliGemma2-3B + AuditDMResolution=448x448, Fine-tuning=per-task2025.12 | 69.2 | |
| PaliGemma2-10BResolution=448x448, Fine-tuning=per-task2025.12 | 68.6 | |
| Llama-3.2-11B-VisionModel Inputs=Question + Image, External Knowledge=Llama-3.2-11B-Vision2025.10 | 67.84 | |
| InternVL3-78BModel Inputs=Question + Image, External Knowledge=InternVL3-78B2025.10 | 67.61 | |
| PaliGemma2-3BResolution=448x448, Fine-tuning=per-task2025.12 | 64.1 | |
| PromptCapEvaluation Protocol=Supervised2023.03 | 58.8 | |
| REVIVEEvaluation Protocol=Supervised2023.03 | 58 | |
| MAILModel Inputs=Question + Image, External Knowledge=Frozen MiniGPT-4 (7B) + ConceptNet2025.10 | 56.69 | |
| RA-VQAEvaluation Protocol=Supervised2023.03 | 54.5 | |
| KAT (Ensemble)Model Inputs=Question + Caption + Object Tags, External Knowledge=Frozen GPT-3 (175B) + Wikidata2025.10 | 54.41 | |
| KATEvaluation Protocol=Supervised2023.03 | 54.4 | |
| REVIVEModel Inputs=Question + Caption + Region Tags, External Knowledge=Frozen GPT-3 (175B) + Wikidata2025.10 | 53.83 | |
| KAT (Single)Model Inputs=Question + Caption + Object Tags, External Knowledge=Frozen GPT-3 (175B) + Wikidata2025.10 | 53.09 | |
| ViperGPTEvaluation Protocol=Zero-shot2023.03 | 51.9 | |
| FlamingoEvaluation Protocol=Zero-shot2023.03 | 50.6 | |
| TRIGEvaluation Protocol=Supervised2023.03 | 50.5 | |
| Pica-FullModel Inputs=Question + Caption + Object Tags, External Knowledge=Frozen GPT-3 (175B)2025.10 | 48 | |
| BLIP-2Evaluation Protocol=Zero-shot2023.03 | 45.9 | |
| MCANModel Inputs=Question + Image, External Knowledge=-2025.10 | 44.65 | |
| PICaEvaluation Protocol=Zero-shot2023.03 | 43.3 | |
| PICA-BaseModel Inputs=Question + Caption + Object Tags, External Knowledge=Frozen GPT-3 (175B)2025.10 | 43.3 | |
| VLC-BERTModel Inputs=Question + Image, External Knowledge=COMET + ConceptNet2025.10 | 43.14 | |
| MAVExModel Inputs=Question + Image, External Knowledge=Wikipedia + ConceptNet + Google Images2025.10 | 41.37 | |
| KrispModel Inputs=Question + Image, External Knowledge=Wikipedia + ConceptNet2025.10 | 38.9 | |
| HCNMNModel Inputs=Question + Image, External Knowledge=WordNet2025.10 | 36.74 | |
| PNP-VQAEvaluation Protocol=Zero-shot2023.03 | 35.9 | |
| ConceptBERTModel Inputs=Question + Image, External Knowledge=ConceptNet2025.10 | 33.66 | |
| MUTAN +ANModel Inputs=Question + Image, External Knowledge=Wikipedia2025.10 | 27.84 | |
| MUTANModel Inputs=Question + Image, External Knowledge=-2025.10 | 26.41 | |
| BAN +ANModel Inputs=Question + Image, External Knowledge=Wikipedia2025.10 | 25.61 | |
| BANModel Inputs=Question + Image, External Knowledge=-2025.10 | 25.17 | |
| Q OnlyModel Inputs=Question + Image, External Knowledge=-2025.10 | 14.93 |