Visual Question Answering on A-OKVQA (val)
79.5AccuracySmoGVLM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SmoGVLMSize=7B, Method=FFT2026.04 | 79.5 | — | — | — | — | |
| LLaVASize=7B, Method=FFT2026.04 | 79.3 | — | — | — | — | |
| SmoGVLM-TinySize=1.3B, Method=FFT2026.04 | 70.7 | — | — | — | — | |
| LLaVA-TinySize=1.3B, Method=FFT2026.04 | 70.4 | — | — | — | — | |
| Qwen2.5-VL-7B + DAPOModel=Qwen2.5-VL-7B, RL Method=DAPO2025.10 | 0.879 | — | — | — | — | |
| Qwen2.5-VL-7B + VPPOModel=Qwen2.5-VL-7B, RL Method=VPPO2025.10 | 0.879 | — | — | — | — | |
| Qwen2.5-VL-7B + GRPOModel=Qwen2.5-VL-7B, RL Method=GRPO2025.10 | 0.874 | — | — | — | — | |
| Qwen3-VL-4b-Instruct2026.03 | 0.852 | — | 0.66 | 0.838 | 0.738 | |
| SNs (mean)Backbone=Qwen3-VL-4b-Instruct2026.03 | 0.852 | 0.8 | 0.66 | 0.837 | 0.738 | |
| SNs (maj. voting)Backbone=Qwen3-VL-4b-Instruct2026.03 | 0.852 | 0.8 | 0.66 | 0.838 | 0.739 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B2025.10 | 0.842 | — | — | — | — | |
| InstructBLIPLLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 0.81 | — | — | — | — | |
| BLIP-2LLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 0.802 | — | — | — | — | |
| InstructBLIPLLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 0.757 | — | — | — | — | |
| Previous SOTAEvaluation Protocol=finetuning2023.05 | 0.732 | — | — | — | — | |
| PromptCap + GPT-3Example retrieval=CLIP (ViT-L/14)2022.11 | 0.732 | — | — | — | — | |
| BLIP-2LLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 0.721 | — | — | — | — | |
| SNs (maj. voting)Backbone=LLaVA-v1.5-7b2026.03 | 0.72 | 0.67 | 0.447 | 0.518 | 0.48 | |
| InstructBLIPLLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 0.64 | — | — | — | — | |
| SNs (mean)Backbone=LLaVA-v1.5-7b2026.03 | 0.614 | 0.67 | 0.377 | 0.837 | 0.52 | |
| GPV-22022.11 | 0.603 | — | — | — | — | |
| GPV-2Evaluation Protocol=Finetune, multimodal=true2023.05 | 0.603 | — | — | — | — | |
| BLIP-2LLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 0.6 | — | — | — | — | |
| PromptCapEvaluation Setting=Supervised, Parameters=175B, Use Extra PLM?=false, With extra V-L Pre-training?=true2023.05 | 0.58 | — | — | — | — | |
| BLIP-2LLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 0.576 | — | — | — | — | |
| InstructBLIPLLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 0.571 | — | — | — | — | |
| CLIPCapEvaluation Protocol=Finetune, multimodal=true2023.05 | 0.5693 | — | — | — | — | |
| Previous SOTAEvaluation Protocol=finetuning2023.05 | 0.563 | — | — | — | — | |
| PromptCap + GPT-3Example retrieval=CLIP (ViT-L/14)2022.11 | 0.563 | — | — | — | — | |
| LLaVA-v1.5-7b2026.03 | 0.548 | — | 0.349 | 0.939 | 0.509 | |
| Text-Davinci-003+ICEvaluation Protocol=Zero-Shot Prompting, multimodal=true, Image Captioning=true2023.05 | 0.5451 | — | — | — | — | |
| KRISP2022.11 | 0.519 | — | — | — | — | |
| KRISPEvaluation Protocol=Finetune, multimodal=true2023.05 | 0.519 | — | — | — | — | |
| LXMERT2022.11 | 0.514 | — | — | — | — | |
| LXMERTEvaluation Protocol=Finetune, multimodal=true2023.05 | 0.514 | — | — | — | — | |
| ViLBERT2022.11 | 0.491 | — | — | — | — | |
| ViLBERTEvaluation Protocol=Finetune, multimodal=true2023.05 | 0.491 | — | — | — | — | |
| Pythia2022.11 | 0.49 | — | — | — | — | |
| PythiaEvaluation Protocol=Finetune, multimodal=true2023.05 | 0.49 | — | — | — | — | |
| GPV-2Evaluation Protocol=Fine-Tuned2022.12 | 0.486 | — | — | — | — | |
| GPV-22022.11 | 0.486 | — | — | — | — | |
| ChatCaptionerEvaluation Protocol=Zero-Shot Prompting, multimodal=true2023.05 | 0.4741 | — | — | — | — | |
| Text-Davinci-003Evaluation Protocol=Few-Shot In-Context Learning (ICL), Shots=22023.05 | 0.4498 | — | — | — | — | |
| BLIP2flanT5xlEvaluation Protocol=Zero-Shot Prompting, multimodal=true2023.05 | 0.448 | — | — | — | — | |
| ClipCap2022.11 | 0.44 | — | — | — | — | |
| Text-Davinci-003Evaluation Protocol=Zero-Shot Prompting2023.05 | 0.4379 | — | — | — | — | |
| ChatGPTEvaluation Protocol=Zero-Shot Prompting2023.05 | 0.433 | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=175B2022.12 | 0.429 | — | — | — | — | |
| Img2LLM 175BEnd-to-End Training=false, Shot Number=02022.12 | 0.429 | — | — | — | — | |
| OFAlargeEvaluation Protocol=Zero-Shot Prompting, multimodal=true2023.05 | 0.4122 | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=66B2022.12 | 0.387 | — | — | — | — | |
| Img2LLM 66BEnd-to-End Training=false, Shot Number=02022.12 | 0.387 | — | — | — | — | |
| BLIPEvaluation Setting=Supervised, Parameters=226M, Use Extra PLM?=false, With extra V-L Pre-training?=false2023.05 | 0.385 | — | — | — | — | |
| LAMOC_11BEvaluation Setting=Zero-shot, Parameters=11.4B, Use Extra PLM?=true, With extra V-L Pre-training?=false2023.05 | 0.379 | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=30B2022.12 | 0.369 | — | — | — | — | |
| Img2LLM 30BEnd-to-End Training=false, Shot Number=02022.12 | 0.369 | — | — | — | — | |
| PNP-VQA_11BEvaluation Setting=Zero-shot, Parameters=11.9B, Use Extra PLM?=true, With extra V-L Pre-training?=false2023.05 | 0.36 | — | — | — | — | |
| GITlargeEvaluation Protocol=Zero-Shot Prompting, multimodal=true2023.05 | 0.3593 | — | — | — | — | |
| PNP-VQA_3BEvaluation Setting=Zero-shot, Parameters=3.9B, Use Extra PLM?=true, With extra V-L Pre-training?=false2023.05 | 0.354 | — | — | — | — | |
| GPT3Evaluation Protocol=Zero-Shot Prompting2023.05 | 0.3507 | — | — | — | — | |
| S³CEvaluation Protocol=unfiltered, Semi-supervised Learning=true2023.09 | 0.342 | — | — | — | — | |
| KRISPEvaluation Protocol=Fine-Tuned2022.12 | 0.337 | — | — | — | — | |
| KRISP2022.11 | 0.337 | — | — | — | — | |
| KRISPEvaluation Protocol=unfiltered2023.09 | 0.337 | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=6.7B2022.12 | 0.333 | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=13B2022.12 | 0.333 | — | — | — | — | |
| Img2LLM 6.7BEnd-to-End Training=false, Shot Number=02022.12 | 0.333 | — | — | — | — | |
| Img2LLM 13BEnd-to-End Training=false, Shot Number=02022.12 | 0.333 | — | — | — | — | |
| Img2Prompt_6.7BEvaluation Setting=Zero-shot, Parameters=8.3B, Use Extra PLM?=true, With extra V-L Pre-training?=false2023.05 | 0.333 | — | — | — | — | |
| Img2Prompt_13BEvaluation Setting=Zero-shot, Parameters=14.6B, Use Extra PLM?=true, With extra V-L Pre-training?=false2023.05 | 0.333 | — | — | — | — | |
| S³C*Evaluation Protocol=unfiltered, Semi-supervised Learning=false2023.09 | 0.33 | — | — | — | — | |
| BERTEvaluation Protocol=Finetune2023.05 | 0.3293 | — | — | — | — | |
| NLX-GPTEvaluation Protocol=unfiltered2023.09 | 0.327 | — | — | — | — | |
| ClipcapEvaluation Protocol=unfiltered2023.09 | 0.308 | — | — | — | — | |
| LXMERTEvaluation Protocol=Fine-Tuned2022.12 | 0.307 | — | — | — | — | |
| LXMERT2022.11 | 0.307 | — | — | — | — | |
| LXMERTEvaluation Protocol=unfiltered2023.09 | 0.307 | — | — | — | — | |
| Most CommonMode=Most Common2023.05 | 0.307 | — | — | — | — | |
| ViLBERTEvaluation Protocol=Fine-Tuned2022.12 | 0.306 | — | — | — | — | |
| ViLBERT2022.11 | 0.306 | — | — | — | — | |
| ViLBERTEvaluation Protocol=unfiltered2023.09 | 0.306 | — | — | — | — | |
| e-UGEvaluation Protocol=unfiltered2023.09 | 0.305 | — | — | — | — | |
| RandomMode=Random2023.05 | 0.267 | — | — | — | — | |
| PythiaEvaluation Protocol=Fine-Tuned2022.12 | 0.252 | — | — | — | — | |
| Pythia2022.11 | 0.252 | — | — | — | — | |
| NLSOM{G,B,O,M}Evaluation Protocol=Zero-Shot Prompting, multimodal=true, rounds=10, Components=Text-Davinci-003 + BLIP2flanT5xl + OFAlarge + mPLUGlarge2023.05 | 0.237 | — | — | — | — | |
| NLSOM{G,B,O}Evaluation Protocol=Zero-Shot Prompting, multimodal=true, rounds=10, Components=Text-Davinci-003 + BLIP2flanT5xl + OFAlarge2023.05 | 0.194 | — | — | — | — | |
| ClipCap Rel->GPT175BEnd-to-End Training=false, Shot Number=102022.12 | 0.181 | — | — | — | — | |
| ClipCap2022.11 | 0.181 | — | — | — | — | |
| ClipCap Cap->GPT175BEnd-to-End Training=false, Shot Number=102022.12 | 0.166 | — | — | — | — | |
| NLSOM{G,O}Evaluation Protocol=Zero-Shot Prompting, multimodal=true, rounds=10, Components=Text-Davinci-003 + OFAlarge2023.05 | 0.139 | — | — | — | — | |
| NLSOM{G,B}Evaluation Protocol=Zero-Shot Prompting, multimodal=true, rounds=10, Components=Text-Davinci-003 + BLIP2flanT5xl2023.05 | 0.091 | — | — | — | — |