Visual Question Answering on VQA 2.0 (test-dev)
86.5AccuracyMolmo-72B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Molmo-72BNumber of crops=36, Prompt=vqa2:2024.09 | 86.5 | — | — | — | — | — | — | — | |
| PaLI-XParameters=55B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 86 | — | — | — | — | — | — | — | |
| InternVL2-Llama-3-76B2024.09 | 85.6 | — | — | — | — | — | — | — | |
| Molmo-7B-DNumber of crops=36, Prompt=vqa2:2024.09 | 85.6 | — | — | — | — | — | — | — | |
| Molmo-7B-ONumber of crops=36, Prompt=vqa2:2024.09 | 85.3 | — | — | — | — | — | — | — | |
| LLaVA OneVision-72Bdistilled=true2024.09 | 85.2 | — | — | — | — | — | — | — | |
| PaLIParameters=17B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 84.3 | — | — | — | — | — | — | — | |
| PaLIVocabulary Setting=open-ended generation, Model Size=17B2023.03 | 84.3 | — | — | — | — | — | — | — | |
| BEIT-3Parameters=1.9B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 84.2 | — | — | — | — | — | — | — | |
| BEiT-3Vocabulary Setting=closed-vocabulary2023.03 | 84.2 | — | — | — | — | — | — | — | |
| BEIT-32022.08 | 84.19 | — | — | — | — | — | — | — | |
| LLaVA OneVision-7Bdistilled=true2024.09 | 84 | — | — | — | — | — | — | — | |
| MolmoE-1BNumber of crops=36, Prompt=vqa2:2024.09 | 83.9 | — | — | — | — | — | — | — | |
| Cambrian-1-34Bdistilled=true2024.09 | 83.8 | — | — | — | — | — | — | — | |
| Qwen2-VL-7B2024.09 | 82.9 | — | — | — | — | — | — | — | |
| VLMO-Large++# Pretrain Images=1.0B2021.11 | 82.88 | — | — | — | — | — | — | — | |
| CoCa2022.08 | 82.3 | — | — | — | — | — | — | — | |
| CoCaVocabulary Setting=closed-vocabulary2023.03 | 82.3 | — | — | — | — | — | — | — | |
| BLIP-2Vocabulary Setting=open-ended generation, Model Size=7B2023.03 | 82.3 | — | — | — | — | — | — | — | |
| OFA2022.08 | 82 | — | — | — | — | — | — | — | |
| Flamingo2022.08 | 82 | — | — | — | — | — | — | — | |
| FlamingoParameters=80B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 82 | — | — | — | — | — | — | — | |
| OFAVocabulary Setting=closed-vocabulary2023.03 | 82 | — | — | — | — | — | — | — | |
| FlamingoVocabulary Setting=open-ended generation, Model Size=80B2023.03 | 82 | — | — | — | — | — | — | — | |
| Qwen2-VL-72B2024.09 | 81.9 | — | — | — | — | — | — | — | |
| GIT2Parameters=5.1B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 81.74 | — | — | — | — | — | — | — | |
| xGen-MM-interleave-4Bdistilled=true2024.09 | 81.5 | — | — | — | — | — | — | — | |
| mPLUG_ViT-LData=14M2022.05 | 81.27 | — | — | — | — | — | — | — | |
| Cambrian-1-8Bdistilled=true2024.09 | 81.2 | — | — | — | — | — | — | — | |
| mPLUG-2#PT Data=17M2023.02 | 81.11 | — | — | — | — | — | — | — | |
| PaLIVocabulary Setting=open-ended generation, Model Size=15B2023.03 | 80.8 | — | — | — | — | — | — | — | |
| MaMMUTVocabulary Setting=open-ended generation, Model Size=2B2023.03 | 80.7 | — | — | — | — | — | — | — | |
| OFA Large#PT Data=18M2023.02 | 80.3 | — | — | — | — | — | — | — | |
| FlorenceVocabulary Setting=closed-vocabulary2023.03 | 80.2 | — | — | — | — | — | — | — | |
| Gemini 1.5 Pro2024.09 | 80.2 | — | — | — | — | — | — | — | |
| Pixtral-12B2024.09 | 80.2 | — | — | — | — | — | — | — | |
| Florence-Huge# Pretrain Images=900M2021.11 | 80.16 | — | — | — | — | — | — | — | |
| Florence2022.08 | 80.16 | — | — | — | — | — | — | — | |
| Florence#PT Data=0.9B2023.02 | 80.16 | — | — | — | — | — | — | — | |
| FlorenceData=0.9B2022.05 | 80.16 | — | — | — | — | — | — | — | |
| Gemini 1.5 Flash2024.09 | 80.1 | — | — | — | — | — | — | — | |
| SimVLM-Huge# Pretrain Images=1.8B2021.11 | 80.03 | — | — | — | — | — | — | — | |
| SimVLM2022.08 | 80.03 | — | — | — | — | — | — | — | |
| SimVLM#PT Data=1.8B2023.02 | 80.03 | — | — | — | — | — | — | — | |
| SimVLMData=1.8B2022.05 | 80.03 | — | — | — | — | — | — | — | |
| PaLM-EParameters=562B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 80 | — | — | — | — | — | — | — | |
| SimVLMVocabulary Setting=closed-vocabulary2023.03 | 80 | — | — | — | — | — | — | — | |
| LLaVA-1.5-13B2024.09 | 80 | — | — | — | — | — | — | — | |
| VLMO-Large# Pretrain Images=4M, Model Size=Large2021.11 | 79.94 | — | — | — | — | — | — | — | |
| VLMo2023.02 | 79.94 | — | — | — | — | — | — | — | |
| VLMo2022.05 | 79.94 | — | — | — | — | — | — | — | |
| OFAData=18M2022.05 | 79.87 | — | — | — | — | — | — | — | |
| mPLUG_ViT-BData=14M2022.05 | 79.79 | — | — | — | — | — | — | — | |
| SimVLM-Large# Pretrain Images=1.8B2021.11 | 79.32 | — | — | — | — | — | — | — | |
| PaLIVocabulary Setting=open-ended generation, Model Size=3B2023.03 | 79.3 | — | — | — | — | — | — | — | |
| mPLUG-2Base#PT Data=17M2023.02 | 79.27 | — | — | — | — | — | — | — | |
| SCLPre-training image scale=<10M images2022.11 | 78.72 | — | — | — | — | — | — | — | |
| GPT-4o-05132024.09 | 78.7 | — | — | — | — | — | — | — | |
| GITVocabulary Setting=open-ended generation2023.03 | 78.6 | — | — | — | — | — | — | — | |
| GIT#PT Data=0.8B2023.02 | 78.56 | — | — | — | — | — | — | — | |
| LLaVA-1.5-7B2024.09 | 78.5 | — | — | — | — | — | — | — | |
| OmniVLVocabulary Setting=closed-vocabulary2023.03 | 78.3 | — | — | — | — | — | — | — | |
| BLIP2022.08 | 78.25 | — | — | — | — | — | — | — | |
| BLIPPre-training image scale=>10M images2022.11 | 78.25 | — | — | — | — | — | — | — | |
| BLIPData=129M2022.05 | 78.25 | — | — | — | — | — | — | — | |
| BLIP#PT=129M2022.06 | 78.24 | — | — | — | — | — | — | — | |
| BLIPVocabulary Setting=open-ended generation2023.03 | 78.2 | — | — | — | — | — | — | — | |
| Llama-3.2V-90B-Instruct2024.09 | 78.1 | — | — | — | — | — | — | — | |
| MAPPre-training dataset=<10M images, Model size=Base size2022.10 | 78.03 | — | — | — | — | — | — | — | |
| OFAPre-training image scale=>10M images2022.11 | 78 | — | — | — | — | — | — | — | |
| mPLUG_ViT-BData=4M2022.05 | 77.94 | — | — | — | — | — | — | — | |
| Unified-IOVocabulary Setting=closed-vocabulary2023.03 | 77.9 | — | — | — | — | — | — | — | |
| SimVLMPre-training image scale=>10M images2022.11 | 77.87 | — | — | — | — | — | — | — | |
| SimVLM-BasePre-training dataset=>10M images, Model size=Base size2022.10 | 77.87 | — | — | — | — | — | — | — | |
| METERVocabulary Setting=closed-vocabulary2023.03 | 77.7 | — | — | — | — | — | — | — | |
| METERPre-training image scale=<10M images2022.11 | 77.68 | — | — | — | — | — | — | — | |
| METERData=4M2022.05 | 77.68 | — | — | — | — | — | — | — | |
| METERPre-training dataset=<10M images, Model size=Base size2022.10 | 77.68 | — | — | — | — | — | — | — | |
| BLIP#PT=14M2022.06 | 77.54 | — | — | — | — | — | — | — | |
| BLIP#PT Data=14M2023.02 | 77.54 | — | — | — | — | — | — | — | |
| Uncompressed2024.03 | 77.4 | — | — | — | — | — | — | 186.1 | |
| GPT-4V2024.09 | 77.2 | — | — | — | — | — | — | — | |
| CCLM_basePre-training Data=4M2022.06 | 77.17 | — | — | — | — | — | — | — | |
| X-VLMclip2022.10 | 76.92 | — | — | — | — | — | — | — | |
| MADTPReduce Ratio=0.52024.03 | 76.8 | — | — | — | — | — | — | 79.4 | |
| InternVL2-8B2024.09 | 76.7 | — | — | — | — | — | — | — | |
| VLMO-Base# Pretrain Images=4M, Model Size=Base2021.11 | 76.64 | — | — | — | — | — | — | — | |
| VLMoPre-training image scale=<10M images2022.11 | 76.64 | — | — | — | — | — | — | — | |
| VLMo-BasePre-training dataset=<10M images, Model size=Base size2022.10 | 76.64 | — | — | — | — | — | — | — | |
| OSCAR+ w/ VinVLModel Size=Large2021.01 | 76.52 | — | — | — | — | — | — | — | |
| VinVL-Large# Pretrain Images=5.7M2021.11 | 76.52 | — | — | — | — | — | — | — | |
| VinVL2022.08 | 76.52 | — | — | — | — | — | — | — | |
| CLIP-ViLpVisual Encoder=CLIP-Res50x4, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 76.48 | — | — | — | — | — | — | — | |
| CLIP-ViLData=4M2022.05 | 76.48 | — | — | — | — | — | — | — | |
| UPopReduce Ratio=0.52024.03 | 76.3 | — | — | — | — | — | — | 109.4 | |
| MADTPReduce Ratio=0.752024.03 | 76.3 | — | — | — | — | — | — | 61.6 | |
| PaliGemma-mix-3B2024.09 | 76.3 | — | — | — | — | — | — | — | |
| EfficientVLM2022.10 | 76.2 | — | — | — | — | — | — | — | |
| Knowledge-CLIPmode=Fine-tuning2022.10 | 76.11 | — | — | — | — | — | — | — | |
| InterBERTModel Size=Large, Ensemble=true2021.01 | 75.95 | — | — | — | — | — | — | — |