Visual Question Answering on VQA v2 (val)
95.06AccuracyOracle
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Oracle2026.03 | 95.06 | — | — | — | — | — | — | — | — | — | |
| LCStraining_mode=learned, evaluation_protocol=5-fold cross-validation2026.03 | 87.38 | — | — | — | — | — | — | — | — | — | |
| MolmoPoint-8BAccess=MolmoPoint2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | |
| HFV-sharptraining_mode=training-free2026.03 | 87.19 | — | — | — | — | — | — | — | — | — | |
| FAAR-learntraining_mode=learned, evaluation_protocol=5-fold cross-validation2026.03 | 87.08 | — | — | — | — | — | — | — | — | — | |
| Molmo2-8BAccess=Fully open2026.03 | 87 | — | — | — | — | — | — | — | — | — | |
| QualRCCVrho=0.4, gamma=1, training_mode=training-free2026.03 | 86.87 | — | — | — | — | — | — | — | — | — | |
| RCCVrho=0.4, training_mode=training-free2026.03 | 86.8 | — | — | — | — | — | — | — | — | — | |
| Calibrated votepool_size=17 models, families=8 families2026.03 | 86.7 | — | — | — | — | — | — | — | — | — | |
| Molmo2-4BAccess=Fully open2026.03 | 86.6 | — | — | — | — | — | — | — | — | — | |
| MolmoPoint-8B-O-7BAccess=Fully open2026.03 | 86.6 | — | — | — | — | — | — | — | — | — | |
| HFVtraining_mode=training-free2026.03 | 86.57 | — | — | — | — | — | — | — | — | — | |
| Single bestpool_size=17 models, families=8 families2026.03 | 86.29 | — | — | — | — | — | — | — | — | — | |
| Majority votepool_size=17 models, families=8 families2026.03 | 86.25 | — | — | — | — | — | — | — | — | — | |
| Specialists SOTAsModel Category=Specialists2024.03 | 86.1 | — | — | — | — | — | — | — | — | — | |
| PLM-8BAccess=Fully open2026.03 | 85.6 | — | — | — | — | — | — | — | — | — | |
| PLM-3BAccess=Fully open2026.03 | 84.4 | — | — | — | — | — | — | — | — | — | |
| Eagle2.5-8BAccess=Open weights only2026.03 | 82.4 | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BAccess=Open weights only2026.03 | 82.3 | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4BAccess=Open weights only2026.03 | 81.7 | — | — | — | — | — | — | — | — | — | |
| LLaVA-NEXT-13B + Visual Lazy AttentionBase Model=LLaVA-NEXT-13B, Optimization Method=Visual Lazy Attention2026.02 | 80.85 | — | — | — | — | — | — | — | — | — | |
| GPT-5Access=API call only2026.03 | 79.7 | — | — | — | — | — | — | — | — | — | |
| MLCDLLM=Qwen2-72B, Vision Tower=ViT-L/14, Training Data=LAION-400M + COYO-700M2024.07 | 79.51 | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8BAccess=Open weights only2026.03 | 79.5 | — | — | — | — | — | — | — | — | — | |
| CLIPLLM=Qwen2-72B, Vision Tower=ViT-L/142024.07 | 79.47 | — | — | — | — | — | — | — | — | — | |
| Keye-VL-1.5-8BAccess=Open weights only2026.03 | 79.3 | — | — | — | — | — | — | — | — | — | |
| FullTraining Data Fraction=100%2026.05 | 79.1 | — | — | — | — | — | — | — | — | — | |
| Qwen-VL + SC-TuneModel Category=Generalists, SC-Tune=true2024.03 | 79 | — | — | — | — | — | — | — | — | — | |
| TPCLFix↑Base=LXMERT, Year=2024, CL=Main technique2024.11 | 78.42 | — | — | — | — | — | — | — | — | — | |
| MLCDLLM=Qwen2-7B, Vision Tower=ViT-L/14, Training Data=LAION-400M + COYO-700M2024.07 | 78.32 | — | — | — | — | — | — | — | — | — | |
| Qwen-VLModel Category=Generalists, SC-Tune=false2024.03 | 78.2 | — | — | — | — | — | — | — | — | — | |
| LocVLM-LLLM=7B, Visual Scale=336, Zero-Shot=false2024.04 | 78.2 | — | — | — | — | — | — | — | — | — | |
| LLaVA-v1.5LLM=7B, Visual Scale=336, Zero-Shot=false2024.04 | 78.1 | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-4BAccess=Open weights only2026.03 | 78.1 | — | — | — | — | — | — | — | — | — | |
| CLIPLLM=Qwen2-7B, Vision Tower=ViT-L/142024.07 | 77.99 | — | — | — | — | — | — | — | — | — | |
| ShikraModel Category=Generalists2024.03 | 77.4 | — | — | — | — | — | — | — | — | — | |
| COIDOTraining Data Fraction=20%2026.05 | 77.2 | — | — | — | — | — | — | — | — | — | |
| Claude Sonnet 4.5Access=API call only2026.03 | 77 | — | — | — | — | — | — | — | — | — | |
| MAGICTraining Data Fraction=20%2026.05 | 76.8 | — | — | — | — | — | — | — | — | — | |
| LLaVA-v1.5-7BBase Model=LLaVA-v1.5-7B, Optimization Method=None2026.02 | 76.64 | — | — | — | — | — | — | — | — | — | |
| COINCIDETraining Data Fraction=20%2026.05 | 76.5 | — | — | — | — | — | — | — | — | — | |
| LLaVA-v1.5-7B + Visual Lazy AttentionBase Model=LLaVA-v1.5-7B, Optimization Method=Visual Lazy Attention2026.02 | 76.38 | — | — | — | — | — | — | — | — | — | |
| ICONSTraining Data Fraction=20%2026.05 | 76.3 | — | — | — | — | — | — | — | — | — | |
| EL2NTraining Data Fraction=20%2026.05 | 76.2 | — | — | — | — | — | — | — | — | — | |
| TPCLDyn↑Base=LXMERT, Year=2024, CL=Main technique2024.11 | 75.83 | — | — | — | — | — | — | — | — | — | |
| PerplexityTraining Data Fraction=20%2026.05 | 75.8 | — | — | — | — | — | — | — | — | — | |
| RandomTraining Data Fraction=20%2026.05 | 75.7 | — | — | — | — | — | — | — | — | — | |
| ShikraLLM=7B, Visual Scale=224, Zero-Shot=false2024.04 | 75.3 | — | — | — | — | — | — | — | — | — | |
| RDSTraining Data Fraction=20%2026.05 | 75.1 | — | — | — | — | — | — | — | — | — | |
| SIMPLEAUGBase=LXMERT, Year=2021, Data Augmentation=Main technique2024.11 | 74.98 | — | — | — | — | — | — | — | — | — | |
| Self-SepTraining Data Fraction=20%2026.05 | 74.9 | — | — | — | — | — | — | — | — | — | |
| SemDeDupTraining Data Fraction=20%2026.05 | 74.2 | — | — | — | — | — | — | — | — | — | |
| Self-FilterTraining Data Fraction=20%2026.05 | 73.7 | — | — | — | — | — | — | — | — | — | |
| MiniGPT-v2 + SC-TuneModel Category=Generalists, SC-Tune=true2024.03 | 73.4 | — | — | — | — | — | — | — | — | — | |
| CLIP-ScoreTraining Data Fraction=20%2026.05 | 73.4 | — | — | — | — | — | — | — | — | — | |
| LXMERTBase=-, Year=20192024.11 | 73.06 | — | — | — | — | — | — | — | — | — | |
| D2-PruningTraining Data Fraction=20%2026.05 | 73 | — | — | — | — | — | — | — | — | — | |
| SRLLM=not specified, Visual Scale=not specified, Zero-Shot=false2024.04 | 72.9 | — | — | — | — | — | — | — | — | — | |
| MiniGPT-v2Model Category=Generalists, SC-Tune=false2024.03 | 72.6 | — | — | — | — | — | — | — | — | — | |
| GPT-5 miniAccess=API call only2026.03 | 72.1 | — | — | — | — | — | — | — | — | — | |
| Gemini 2.5 FlashAccess=API call only2026.03 | 69.4 | — | — | — | — | — | — | — | — | — | |
| GLM-4.1V-9BAccess=Open weights only2026.03 | 68.3 | — | — | — | — | — | — | — | — | — | |
| VilBERTpretrained=Conceptual Caption, fine-tuned=VQA v22021.04 | 67.77 | — | — | — | — | — | — | — | — | — | |
| Gemini 2.5 ProAccess=API call only2026.03 | 67.1 | — | — | — | — | — | — | — | — | — | |
| DGGBase=UpDn, Year=2023, Data Augmentation=Main technique2024.11 | 65.54 | — | — | — | — | — | — | — | — | — | |
| BLIP-2 ViT-g FlanT5XXL#Trainable Params=108M, #Total Params=12.1B, zero-shot=true2023.01 | 65.2 | — | — | — | — | — | — | — | — | — | |
| BLIP-2 ViT-G FlanT5XXLGPU hours=684.0, training data=121.6M2023.05 | 65.2 | — | — | — | — | — | — | — | — | — | |
| VPGTrans (BLIP-2 ViT-G FlanT5XL -> XXL)GPU hours=32.4, training data=5.3M2023.05 | 65.2 | — | — | — | — | — | — | — | — | — | |
| BLIP-2Model Category=Generalists2024.03 | 65 | — | — | — | — | — | — | — | — | — | |
| D-VQABase=LXMERT, Year=2021, Ensemble Learning=Main technique2024.11 | 64.96 | — | — | — | — | — | — | — | — | — | |
| FAN-VQABase=UpDn, Year=2024, Ensemble Learning=Secondary technique, Data Augmentation=Main technique2024.11 | 64.92 | — | — | — | — | — | — | — | — | — | |
| MiniCPM-V-4.5-8BAccess=Open weights only2026.03 | 64.1 | — | — | — | — | — | — | — | — | — | |
| GVQE*Base=UpDn, extra_annotations=true2021.07 | 64.04 | — | — | — | — | — | — | — | — | — | |
| BLOCK2021.04 | 63.89 | — | — | — | — | — | — | — | — | — | |
| UpDnBase=N/A2021.07 | 63.79 | — | — | — | — | — | — | 80.94 | 42.51 | 55.78 | |
| CF-VQA(Sum)Base=UpDn2021.07 | 63.65 | — | — | — | — | — | — | 82.63 | 44.01 | 54.38 | |
| LfFBase architecture=UpDown [3]2021.04 | 63.57 | — | — | — | — | — | — | — | — | — | |
| CF-VQABase=UpDn, Year=2021, Ensemble Learning=Main technique2024.11 | 63.54 | — | — | — | — | — | — | — | — | — | |
| UpDown2021.04 | 63.52 | — | — | — | — | — | — | — | — | — | |
| UpDnBase=-, Year=20182024.11 | 63.48 | — | — | — | — | — | — | — | — | — | |
| HINT*Base=UpDn, extra_annotations=true2021.07 | 63.38 | — | — | — | — | — | — | 81.18 | 42.14 | 55.66 | |
| HINTBase=UpDn, Year=2019, Answer Re-Ranking=Main technique2024.11 | 63.38 | — | — | — | — | — | — | — | — | — | |
| RandImgBase architecture=UpDown [3]2021.04 | 63.34 | — | — | — | — | — | — | — | — | — | |
| LMBase=UpDn2021.07 | 63.26 | — | — | — | — | — | — | 81.16 | 42.22 | 55.22 | |
| AttAlignBase=UpDn, Year=2019, Answer Re-Ranking=Main technique2024.11 | 63.24 | — | — | — | — | — | — | — | — | — | |
| GVQE*Base=S-MRL, extra_annotations=true2021.07 | 63.18 | — | — | — | — | — | — | — | — | — | |
| BLIP-2 ViT-g FlanT5XL#Trainable Params=107M, #Total Params=4.1B, zero-shot=true2023.01 | 63.1 | — | — | — | — | — | — | — | — | — | |
| S-MRLBase=N/A2021.07 | 63.1 | — | — | — | — | — | — | — | — | — | |
| ESRBase architecture=UpDown [3]2021.04 | 62.96 | — | — | — | — | — | — | — | — | — | |
| AdvReg.Base=UpDn2021.07 | 62.75 | — | — | — | — | — | — | 79.84 | 42.35 | 55.16 | |
| BLIP-2 ViT-L FlanT5XL#Trainable Params=103M, #Total Params=3.4B, zero-shot=true2023.01 | 62.6 | — | — | — | — | — | — | — | — | — | |
| MutantBase=UpDn, Year=2020, Data Augmentation=Main technique2024.11 | 62.56 | — | — | — | — | — | — | — | — | — | |
| SCRBase=UpDn, Year=2019, Answer Re-Ranking=Main technique2024.11 | 62.3 | — | — | — | — | — | — | — | — | — | |
| SCR*Base=UpDn, extra_annotations=true2021.07 | 62.2 | — | — | — | — | — | — | 78.8 | 41.6 | 54.4 | |
| RUBiBase architecture=UpDown [3]2021.04 | 61.88 | — | — | — | — | — | — | — | — | — | |
| RUBiBase=S-MRL2021.07 | 61.16 | — | — | — | — | — | — | — | — | — | |
| LMHBase architecture=UpDown [3]2021.04 | 61.15 | — | — | — | — | — | — | — | — | — | |
| LMH + RMFEBase architecture=UpDown [3]2021.04 | 60.96 | — | — | — | — | — | — | — | — | — | |
| CF-VQA(Sum)Base=S-MRL2021.07 | 60.76 | — | — | — | — | — | — | 81.11 | 43.48 | 49.58 | |
| CTRL-ONumber of Parameters=39M, Coupling=true2025.03 | 60.25 | — | — | — | — | — | — | — | — | — |