Visual Question Answering on VQA v2 (test-std)
86.1AccuracyPaLI-X
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| PaLI-XParameters=55B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 86.1 | — | — | — | — | — | |
| PaLI-XNumber of parameters=55B2023.05 | 86.1 | — | — | — | — | — | |
| PaLI-XParameters=55B2023.10 | 86.1 | — | — | — | — | — | |
| PaLI-3Parameters=5B2023.10 | 85.2 | — | — | — | — | — | |
| PaLIEvaluation Protocol=Task-specific finetuned2023.03 | 84.3 | — | — | — | — | — | |
| PaLI-17BVocabulary setting=Open-vocabulary generation2022.09 | 84.3 | — | — | — | — | — | |
| PaLIParameters=17B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 84.3 | — | — | — | — | — | |
| PaLINumber of parameters=17B2023.05 | 84.3 | — | — | — | — | — | |
| PaLIVocabulary Setting=open-ended generation, Model Size=17B2023.03 | 84.3 | — | — | — | — | — | |
| PaLI-17BParameters=17B2023.10 | 84.3 | — | — | — | — | — | |
| PaLIPre-train (# Pairs)=1.6B, Evaluation Setting=Generative2023.03 | 84.3 | — | — | — | — | — | |
| BEIT-3#Trainable Params=1.9B, Model Type=Closed-ended classification2023.01 | 84.03 | — | — | — | — | — | |
| BEiT-3Example=21M2023.06 | 84.03 | — | — | — | — | — | |
| BEiT-3Parameters=1.9B, Vocabulary setting=Closed-vocabulary classification2022.09 | 84 | — | — | — | — | — | |
| BEIT-3Parameters=1.9B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 84 | — | — | — | — | — | |
| BEiT-3Number of parameters=1.9B2023.05 | 84 | — | — | — | — | — | |
| BEIT-32023.05 | 84 | — | — | — | — | — | |
| BEiT-3Vocabulary Setting=closed-vocabulary2023.03 | 84 | — | — | — | — | — | |
| BEIT-3Parameters=1.9B2023.10 | 84 | — | — | — | — | — | |
| BEiT-3# Params=1.9B2023.01 | 84 | — | — | — | — | — | |
| ONE-PEACE2023.05 | 82.5 | — | — | — | — | — | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 82.3 | — | — | — | — | — | |
| CoCaParameters=2.1B, Vocabulary setting=Closed-vocabulary classification2022.09 | 82.3 | — | — | — | — | — | |
| CoCaNumber of parameters=2.1B2023.05 | 82.3 | — | — | — | — | — | |
| CoCa2023.05 | 82.3 | — | — | — | — | — | |
| BLIP-22023.05 | 82.3 | — | — | — | — | — | |
| BLIP-2 ViT-g OPT6.7B#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 82.3 | — | — | — | — | — | |
| CoCa#Trainable Params=2.1B, Model Type=Closed-ended classification2023.01 | 82.3 | — | — | — | — | — | |
| BLIP-2Example=129M2023.06 | 82.3 | — | — | — | — | — | |
| CoCaExample=4.8B2023.06 | 82.3 | — | — | — | — | — | |
| CoCaVocabulary Setting=closed-vocabulary2023.03 | 82.3 | — | — | — | — | — | |
| CoCaParameters=2.1B2023.10 | 82.3 | — | — | — | — | — | |
| CoCa# Params=2.1B2023.01 | 82.3 | — | — | — | — | — | |
| CoCaPre-train (# Pairs)=4.8B, Evaluation Setting=Discriminative2023.03 | 82.3 | — | — | — | — | — | |
| Flamingomode=Fine-tuned2022.04 | 82.1 | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Fine-tuned2022.04 | 82.1 | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Task-specific finetuned2023.03 | 82.1 | — | — | — | — | — | |
| FlamingoParameters=80B, Vocabulary setting=Open-vocabulary generation2022.09 | 82.1 | — | — | — | — | — | |
| FlamingoParameters=80B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 82.1 | — | — | — | — | — | |
| FlamingoNumber of parameters=80B2023.05 | 82.1 | — | — | — | — | — | |
| Flamingo2023.05 | 82.1 | — | — | — | — | — | |
| Flamingo80B#Trainable Params=10.6B, Model Type=Open-ended generation2023.01 | 82.1 | — | — | — | — | — | |
| FlamingoExample=2.3B2023.06 | 82.1 | — | — | — | — | — | |
| FlamingoVocabulary Setting=open-ended generation, Model Size=80B2023.03 | 82.1 | — | — | — | — | — | |
| FlamingoParameters=80B, Shots=322023.10 | 82.1 | — | — | — | — | — | |
| OFA2022.02 | 82 | — | — | — | — | — | |
| OFAParameters=0.9B, Vocabulary setting=Closed-vocabulary classification2022.09 | 82 | — | — | — | — | — | |
| OFANumber of parameters=0.9B2023.05 | 82 | — | — | — | — | — | |
| OFA2023.05 | 82 | — | — | — | — | — | |
| OFA#Trainable Params=930M, Model Type=Open-ended generation2023.01 | 82 | — | — | — | — | — | |
| OFAExample=18M2023.06 | 82 | — | — | — | — | — | |
| OFAVocabulary Setting=closed-vocabulary2023.03 | 82 | — | — | — | — | — | |
| OFAParameters=0.9B2023.10 | 82 | — | — | — | — | — | |
| GIT2Parameters=5.1B, Vocabulary setting=Closed-vocabulary classification2022.09 | 81.92 | — | — | — | — | — | |
| GIT2Parameters=5.1B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 81.92 | — | — | — | — | — | |
| GIT2Number of parameters=5.1B2023.05 | 81.92 | — | — | — | — | — | |
| GIT2Parameters=5.1B2023.10 | 81.92 | — | — | — | — | — | |
| GIT2Example=12.9B2023.06 | 81.9 | — | — | — | — | — | |
| GIT-2Pre-train (# Pairs)=12.9B, Evaluation Setting=Generative2023.03 | 81.9 | — | — | — | — | — | |
| X2-VLM_large# Params=593M, Pre-training Data=More Data2022.11 | 81.8 | — | — | — | — | — | |
| BLIP-2 ViT-g OPT2.7B#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 81.74 | — | — | — | — | — | |
| BLIP-2 ViT-g FlanT5xL#Trainable Params=1.2B, Model Type=Open-ended generation2023.01 | 81.66 | — | — | — | — | — | |
| SotA2022.04 | 81.3 | — | — | — | — | — | |
| Unrestricted SotA (Yan et al.)Evaluation Protocol=Fine-tuned2022.04 | 81.3 | — | — | — | — | — | |
| mPLUG_ViT-LData=14M2022.05 | 81.26 | — | — | — | — | — | |
| mPLUG-2#PT Data=17M2023.02 | 81.13 | — | — | — | — | — | |
| mPLUG-2Example=417M2023.06 | 81.13 | — | — | — | — | — | |
| MaMMUTVocabulary Setting=open-ended generation, Model Size=2B2023.03 | 80.8 | — | — | — | — | — | |
| METER-CoSwinHUGE# Pre-training Images=14M2021.11 | 80.54 | — | — | — | — | — | |
| COSAExample=415M, parameters=1.2B2023.06 | 80.54 | — | — | — | — | — | |
| OFALarge2022.02 | 80.5 | — | — | — | — | — | |
| METERInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.5 | — | — | — | — | — | |
| OFA Large#PT Data=18M2023.02 | 80.5 | — | — | — | — | — | |
| X2-VLM_large# Params=593M, Pre-training Data=4M2022.11 | 80.5 | — | — | — | — | — | |
| OFA_large# Params=472M, Pre-training Data=More Data2022.11 | 80.5 | — | — | — | — | — | |
| FlorenceBackbone=CoSwin-H + RoBERTa, Pre-trained Data=900M image-text pairs2022.02 | 80.4 | — | — | — | — | — | |
| FlorenceInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.4 | — | — | — | — | — | |
| FlorenceEvaluation Protocol=Fine-tuned2022.04 | 80.4 | — | — | — | — | — | |
| Restricted SotA+ (Florence)Evaluation Protocol=Fine-tuned2022.04 | 80.4 | — | — | — | — | — | |
| FlorenceVocabulary Setting=closed-vocabulary2023.03 | 80.4 | — | — | — | — | — | |
| X-FMbase# Params=327M, Training Data=More Data2023.01 | 80.4 | — | — | — | — | — | |
| Florence2021.11 | 80.36 | — | — | — | — | — | |
| Florence#PT Data=0.9B2023.02 | 80.36 | — | — | — | — | — | |
| FlorenceExample=900M2023.06 | 80.36 | — | — | — | — | — | |
| FlorenceData=0.9B2022.05 | 80.36 | — | — | — | — | — | |
| SimVLM2021.11 | 80.34 | — | — | — | — | — | |
| SimVLM HUGE# Pre-training Images=1.8B2021.11 | 80.34 | — | — | — | — | — | |
| SimVLM_HUGEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 80.34 | — | — | — | — | — | |
| SimVLMModel Size=Huge2021.08 | 80.34 | — | — | — | — | — | |
| SimVLM-HPretrain Images=1.8B2022.06 | 80.34 | — | — | — | — | — | |
| SimVLM#PT Data=1.8B2023.02 | 80.34 | — | — | — | — | — | |
| SimVLMVocabulary setting=Closed-vocabulary classification2022.09 | 80.34 | — | — | — | — | — | |
| SimVLM2023.05 | 80.34 | — | — | — | — | — | |
| SimVLM#Trainable Params=~1.4B, Model Type=Closed-ended classification2023.01 | 80.34 | — | — | — | — | — | |
| SimVLMExample=1.8B2023.06 | 80.34 | — | — | — | — | — | |
| SimVLMData=1.8B2022.05 | 80.34 | — | — | — | — | — | |
| SimVLM2023.10 | 80.34 | — | — | — | — | — | |
| SimVLMModel Size=Huge, Pre-trained Data=1.8B image-text pairs, Backbone=ViT-Huge2022.02 | 80.3 | — | — | — | — | — | |
| SimVLMInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.3 | — | — | — | — | — | |
| SimVLMEvaluation Protocol=Fine-tuned2022.04 | 80.3 | — | — | — | — | — |