Visual Question Answering on VQA v2 (test-dev)
87.66Overall AccuracyLCS
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LCSTraining set=full minival set, Model pool size=12 models, Family count=5 families2026.03 | 87.66 | 97.35 | 80.77 | 80.88 | — | — | |
| PaLI-XNumber of parameters=55B2023.05 | 86 | — | — | — | — | — | |
| PaLI-XParameters=55B2023.10 | 86 | — | — | — | — | — | |
| PaLI-XBase model=PaLI-X2023.12 | 86 | — | — | — | — | — | |
| SOTA [7]2023.12 | 86 | — | — | — | — | — | |
| SMoLA-PaLI-X_FT (Specialist)Base model=PaLI-X, Per-task LoRA tuning=true, LoRA rank=42023.12 | 85.7 | — | — | — | — | — | |
| PerceptionGPT2023.11 | 85.1 | — | — | — | — | — | |
| PaLI-3Parameters=5B2023.10 | 85 | — | — | — | — | — | |
| Emu2-ChatLLM=LLaMA-33B, Trained during SFT stage=true2023.11 | 84.9 | — | — | — | — | — | |
| PaLIEvaluation Protocol=Task-specific finetuned2023.03 | 84.3 | — | — | — | — | — | |
| PaLI-17BVocabulary setting=Open-vocabulary generation2022.09 | 84.3 | — | — | — | — | — | |
| PaLINumber of parameters=17B2023.05 | 84.3 | — | — | — | — | — | |
| PaLI-17BParameters=17B2023.10 | 84.3 | — | — | — | — | — | |
| PaLIPre-train (# Pairs)=1.6B, Evaluation Setting=Generative2023.03 | 84.3 | — | — | — | — | — | |
| PaLI-17Btrained on QA datasets=true2024.03 | 84.3 | — | — | — | — | — | |
| PaLI-17Btrained_on_qa_datasets=true2024.03 | 84.3 | — | — | — | — | — | |
| BEiT-3Parameters=1.9B, Vocabulary setting=Closed-vocabulary classification2022.09 | 84.2 | — | — | — | — | — | |
| BEiT-3Number of parameters=1.9B2023.05 | 84.2 | — | — | — | — | — | |
| BEIT-32023.05 | 84.2 | — | — | — | — | — | |
| BEIT-3Parameters=1.9B2023.10 | 84.2 | — | — | — | — | — | |
| BEiT-3# Params=1.9B2023.01 | 84.2 | — | — | — | — | — | |
| BEIT-3trained on QA datasets=true2024.03 | 84.2 | — | — | — | — | — | |
| BEIT-3trained_on_qa_datasets=true2024.03 | 84.2 | — | — | — | — | — | |
| BEIT-3#Trainable Params=1.9B, Model Type=Closed-ended classification2023.01 | 84.19 | — | — | — | — | — | |
| BEIT-32022.08 | 84 | — | — | — | — | — | |
| MM1-ChatSize=30B, # tokens per image=720, Zero-shot=true2024.05 | 83.7 | — | — | — | — | — | |
| LLaVA-NeXTSize=34B, # tokens per image=2880, Zero-shot=true2024.05 | 83.7 | — | — | — | — | — | |
| ITNet-LParams=307M2026.06 | 83.6 | — | — | — | — | — | |
| Shikra2023.11 | 83.3 | — | — | — | — | — | |
| PaLI-15BVocabulary setting=Open-vocabulary generation2022.09 | 82.9 | — | — | — | — | — | |
| Bunny-8BSize=4B < Size < 8B2024.02 | 82.9 | — | — | — | — | — | |
| LLaVA-NeXTSize=13B, # tokens per image=2880, Zero-shot=true2024.05 | 82.8 | — | — | — | — | — | |
| MM1-ChatSize=7B, # tokens per image=720, Zero-shot=true2024.05 | 82.8 | — | — | — | — | — | |
| VILA1.5-13BSize=> 8B2024.02 | 82.8 | — | — | — | — | — | |
| LLaVA-NeXT-13BSize=> 8B2024.02 | 82.8 | — | — | — | — | — | |
| MM1-7B-ChatSize=4B < Size < 8B2024.02 | 82.8 | — | — | — | — | — | |
| VLMOtrained on QA datasets=true2024.03 | 82.8 | — | — | — | — | — | |
| VLMOtrained_on_qa_datasets=true2024.03 | 82.8 | — | — | — | — | — | |
| ONE-PEACE2023.05 | 82.6 | — | — | — | — | — | |
| FMVR-LLaVA#Vision Tokens=28802026.03 | 82.4 | — | — | — | — | — | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 82.3 | — | — | — | — | — | |
| CoCa2022.08 | 82.3 | — | — | — | — | — | |
| CoCaStatus=Fine-tuned SOTA2022.06 | 82.3 | — | — | — | — | — | |
| CoCaParameters=2.1B, Vocabulary setting=Closed-vocabulary classification2022.09 | 82.3 | — | — | — | — | — | |
| CoCaNumber of parameters=2.1B2023.05 | 82.3 | — | — | — | — | — | |
| CoCa2023.05 | 82.3 | — | — | — | — | — | |
| CoCa#Trainable Params=2.1B, Model Type=Closed-ended classification2023.01 | 82.3 | — | — | — | — | — | |
| CoCaParameters=2.1B2023.10 | 82.3 | — | — | — | — | — | |
| CoCa# Params=2.1B2023.01 | 82.3 | — | — | — | — | — | |
| CoCaPre-train (# Pairs)=4.8B, Evaluation Setting=Discriminative2023.03 | 82.3 | — | — | — | — | — | |
| CogVLM-ChatLLM=Vicuna-7B, Trained during SFT stage=true2023.11 | 82.3 | — | — | — | — | — | |
| CoCaVocabulary=Closed2022.05 | 82.3 | — | — | — | — | — | |
| BLIP-22023.05 | 82.2 | — | — | — | — | — | |
| BLIP-2 ViT-g OPT6.7B#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 82.19 | — | — | — | — | — | |
| Bunny-4BSize=Size < 4B2024.02 | 82.1 | — | — | — | — | — | |
| OFA2022.02 | 82 | — | — | — | — | — | |
| Flamingomode=Fine-tuned2022.04 | 82 | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Fine-tuned2022.04 | 82 | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Task-specific finetuned2023.03 | 82 | — | — | — | — | — | |
| OFAParameters=0.9B, Vocabulary setting=Closed-vocabulary classification2022.09 | 82 | — | — | — | — | — | |
| FlamingoParameters=80B, Vocabulary setting=Open-vocabulary generation2022.09 | 82 | — | — | — | — | — | |
| OFANumber of parameters=0.9B2023.05 | 82 | — | — | — | — | — | |
| FlamingoNumber of parameters=80B2023.05 | 82 | — | — | — | — | — | |
| OFA2023.05 | 82 | — | — | — | — | — | |
| Flamingo2023.05 | 82 | — | — | — | — | — | |
| OFA#Trainable Params=930M, Model Type=Open-ended generation2023.01 | 82 | — | — | — | — | — | |
| Flamingo80B#Trainable Params=10.6B, Model Type=Open-ended generation2023.01 | 82 | — | — | — | — | — | |
| OFAParameters=0.9B2023.10 | 82 | — | — | — | — | — | |
| FlamingoParameters=80B, Shots=322023.10 | 82 | — | — | — | — | — | |
| MM1-3B-ChatSize=Size < 4B2024.02 | 82 | — | — | — | — | — | |
| FlamingoVocabulary=Open, Parameters=80B2022.05 | 82 | — | — | — | — | — | |
| X2-VLM_large# Params=593M, Pre-training Data=More Data2022.11 | 81.9 | — | — | — | — | — | |
| LLaVA-NeXT-7BSize=4B < Size < 8B2024.02 | 81.8 | — | — | — | — | — | |
| GIT2Parameters=5.1B, Vocabulary setting=Closed-vocabulary classification2022.09 | 81.74 | — | — | — | — | — | |
| GIT2Number of parameters=5.1B2023.05 | 81.74 | — | — | — | — | — | |
| GIT2Parameters=5.1B2023.10 | 81.74 | — | — | — | — | — | |
| GIT2Vocabulary=Open, Parameters=5.1B2022.05 | 81.74 | — | — | — | — | — | |
| GIT-2Pre-train (# Pairs)=12.9B, Evaluation Setting=Generative2023.03 | 81.7 | — | — | — | — | — | |
| BLIP-2 ViT-g OPT2.7B#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 81.59 | — | — | — | — | — | |
| BLIP-2 ViT-g FlanT5xL#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 81.55 | — | — | — | — | — | |
| Imp-v1.5-4B-Phi3Size=Size < 4B2024.02 | 81.5 | — | — | — | — | — | |
| PaLI-3BVocabulary setting=Open-vocabulary generation2022.09 | 81.4 | — | — | — | — | — | |
| SotA2022.04 | 81.3 | — | — | — | — | — | |
| Unrestricted SotA (Yan et al.)Evaluation Protocol=Fine-tuned2022.04 | 81.3 | — | — | — | — | — | |
| Mipha-3BSize=Size < 4B2024.02 | 81.3 | — | — | — | — | — | |
| LLaVA-NeXT-7B#Vision Tokens=28802026.03 | 81.3 | — | — | — | — | — | |
| mPlugVocabulary=Closed2022.05 | 81.27 | — | — | — | — | — | |
| Idefics2Size=8B, # tokens per image=320, Zero-shot=true2024.05 | 81.2 | — | — | — | — | — | |
| Idefics2Size=4B < Size < 8B2024.02 | 81.2 | — | — | — | — | — | |
| ShareGPT4V-13BSize=> 8B2024.02 | 81 | — | — | — | — | — | |
| Idefics2Size=8B, # tokens per image=64, Zero-shot=true2024.05 | 80.8 | — | — | — | — | — | |
| SPHINX-2kLLM=LLaMA2 13B, Trained during SFT stage=true2023.11 | 80.7 | — | — | — | — | — | |
| LVIS-INSTRUCT4V-13BSize=> 8B2024.02 | 80.7 | — | — | — | — | — | |
| PaLM-E-84BEvaluation Protocol=Task-specific finetuned2023.03 | 80.5 | — | — | — | — | — | |
| X2-VLM_large# Params=593M, Pre-training Data=4M2022.11 | 80.5 | — | — | — | — | — | |
| X-FMbase# Params=327M, Training Data=More Data2023.01 | 80.5 | — | — | — | — | — | |
| X-FMModel size=base-size2023.01 | 80.5 | — | — | — | — | — | |
| LLaVA-1.5 Vicuna-13B w/ EasyGenInstruction Tuning Data (IT)=665K2023.10 | 80.5 | — | — | — | — | — | |
| X2-VLM_base# Params=255M, Pre-training Data=More Data2022.11 | 80.4 | — | — | — | — | — | |
| X2-VLMbase# Params=255M, Training Data=More Data2023.01 | 80.4 | — | — | — | — | — |