Fact-based Visual Question Answering on FVQA (test)
72.61ScoreGemini-3-Pro
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-3-ProWorkflow=Agentic Model, Zero-shot=true2026.07 | 72.61 | — | — | — | |
| GPT-5.2Workflow=Agentic Model, Zero-shot=true2026.07 | 68.78 | — | — | — | |
| GPT-4oWorkflow=Agentic Model, Zero-shot=true2026.07 | 66.34 | — | — | — | |
| Gemini-3-FlashWorkflow=Agentic Model, Zero-shot=true2026.07 | 64.89 | — | — | — | |
| VideoSearcher-8BWorkflow=Agentic Model, Zero-shot=false2026.07 | 64.1 | — | — | — | |
| DeepEyesV2Workflow=Agentic Model, Zero-shot=false2026.07 | 60.6 | — | — | — | |
| Gemini-3-ProWorkflow=Direct Answer, Zero-shot=false2026.07 | 59.22 | — | — | — | |
| MMSearch-R1Workflow=Agentic Model, Zero-shot=false2026.07 | 58.4 | — | — | — | |
| Webwatcher-7BWorkflow=Agentic Model, Zero-shot=false2026.07 | 58.17 | — | — | — | |
| Gemini-3-FlashWorkflow=Direct Answer, Zero-shot=false2026.07 | 56.5 | — | — | — | |
| VideoSearcher-4BWorkflow=Agentic Model, Zero-shot=false2026.07 | 55.17 | — | — | — | |
| Qwen3-VL-32B-InstructWorkflow=Agentic Model, Zero-shot=true2026.07 | 54.28 | — | — | — | |
| Qwen3-VL-8B-InstructWorkflow=Agentic Model, Zero-shot=true2026.07 | 53.61 | — | — | — | |
| Qwen2.5-VL-32B-InstructWorkflow=Agentic Model, Zero-shot=true2026.07 | 52.22 | — | — | — | |
| GPT-5.2Workflow=Direct Answer, Zero-shot=false2026.07 | 50.94 | — | — | — | |
| GPT-4oWorkflow=Direct Answer, Zero-shot=false2026.07 | 48 | — | — | — | |
| Visual-ARFTWorkflow=Agentic Model, Zero-shot=false2026.07 | 41.72 | — | — | — | |
| Qwen2.5-VL-7B-InstructWorkflow=Agentic Model, Zero-shot=true2026.07 | 36 | — | — | — | |
| Qwen3-VL-32B-InstructWorkflow=Direct Answer, Zero-shot=false2026.07 | 32.17 | — | — | — | |
| Qwen3-VL-4B-InstructWorkflow=Agentic Model, Zero-shot=true2026.07 | 31.72 | — | — | — | |
| Qwen2.5-VL-32B-InstructWorkflow=Direct Answer, Zero-shot=false2026.07 | 30.5 | — | — | — | |
| Qwen2.5-VL-7B-InstructWorkflow=Direct Answer, Zero-shot=false2026.07 | 26.28 | — | — | — | |
| Qwen3-VL-4B-InstructWorkflow=Direct Answer, Zero-shot=false2026.07 | 25.33 | — | — | — | |
| Qwen3-VL-8B-InstructWorkflow=Direct Answer, Zero-shot=false2026.07 | 24.22 | — | — | — | |
| FVQA modelQuestion-Query Mapping Strategy=ground truth mapping (gt-QQmaping)2016.06 | — | 65.51 | 72.37 | 73.55 | |
| FVQA modelQuestion-Query Mapping Strategy=top-1 predicted mapping2016.06 | — | 54.79 | 61.41 | 62.22 | |
| FVQA modelQuestion-Query Mapping Strategy=top-3 predicted mapping2016.06 | — | 59.67 | 66.89 | 67.77 | |
| Hie-Question+ImageArchitecture=Hierarchical, Input=Question+Image2016.06 | — | 39.75 | 56.48 | 69.78 | |
| Hie-Question+Image+Pre-VQAArchitecture=Hierarchical, Input=Question+Image, Pre-training=Pre-VQA2016.06 | — | 48.93 | 64.75 | 76.73 | |
| Human2016.06 | — | 82.47 | — | — | |
| LSTM-ImageArchitecture=LSTM, Input=Image2016.06 | — | 26.78 | 44 | 62.86 | |
| LSTM-QuestionArchitecture=LSTM, Input=Question2016.06 | — | 15.82 | 26.45 | 40.99 | |
| LSTM-Question+ImageArchitecture=LSTM, Input=Question+Image2016.06 | — | 29.08 | 44.36 | 61.71 | |
| LSTM-Question+Image+Pre-VQAArchitecture=LSTM, Input=Question+Image, Pre-training=Pre-VQA2016.06 | — | 31.96 | 48.55 | 64.73 | |
| SVM-ImageArchitecture=SVM, Input=Image2016.06 | — | 24.73 | 40.95 | 56.33 | |
| SVM-QuestionArchitecture=SVM, Input=Question2016.06 | — | 17.06 | 29.43 | 44.73 | |
| SVM-Question+ImageArchitecture=SVM, Input=Question+Image2016.06 | — | 25.3 | 41.37 | 56.78 |