Natural Language Visual Reasoning on NLVR2 (test-p)
92.6AccuracyBEIT-3
Evaluation Results
| Method | Links | |
|---|---|---|
| BEIT-32023.05 | 92.6 | |
| BEIT-32022.08 | 92.58 | |
| VLMO-Large++# Pretrain Images=1.0B2021.11 | 89.54 | |
| X2-VLM_large# Params=593M, Pre-training Data=More Data2022.11 | 89.4 | |
| ONE-PEACE2023.05 | 88.3 | |
| X2-VLM_large# Params=593M, Pre-training Data=4M2022.11 | 87.6 | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 87 | |
| CoCa2022.08 | 87 | |
| CoCa2023.05 | 87 | |
| CoCa#Images=3B, Para.=2.1B2022.12 | 87 | |
| X2-VLM_base# Params=255M, Pre-training Data=More Data2022.11 | 87 | |
| VLMoInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 86.9 | |
| VLMo_large# Params=562M, Pre-training Data=4M2022.11 | 86.9 | |
| VLMO-Large# Pretrain Images=4M, Model Size=Large2021.11 | 86.86 | |
| VLMo-LPretrain Images=4M2022.06 | 86.86 | |
| VLMo2022.05 | 86.86 | |
| X2-VLM_base# Params=255M, Pre-training Data=4M2022.11 | 86.1 | |
| FIBER-BPretrain Images=4M2022.06 | 85.52 | |
| SimVLMInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 85.2 | |
| SimVLM2023.05 | 85.2 | |
| SimVLMModel Size=Huge2021.08 | 85.15 | |
| SimVLM-Huge# Pretrain Images=1.8B2021.11 | 85.15 | |
| SimVLM2022.08 | 85.15 | |
| SimVLM-HPretrain Images=1.8B2022.06 | 85.15 | |
| mPLUG2022.05 | 84.95 | |
| SimVLMModel Size=Large2021.08 | 84.84 | |
| SimVLM-Large# Pretrain Images=1.8B2021.11 | 84.84 | |
| SimVLM-LARGE# Pre-train Images=1.8B, Model scale=Large2023.05 | 84.84 | |
| SimVLMSize=Large2022.05 | 84.84 | |
| SimVLM_large# Params=783M, Pre-training Data=More Data2022.11 | 84.8 | |
| X-VLMPre-training Images Scale=16M2021.11 | 84.76 | |
| X-VLMPretrain Images=4M2022.06 | 84.76 | |
| X-VLM2023.05 | 84.76 | |
| SCLPre-training image scale=<10M images2022.11 | 84.27 | |
| X-VLMPre-training Images Scale=4M2021.11 | 84.21 | |
| SimVLM-BASE# Pre-train Images=1.8B, Model scale=Base2023.05 | 84.15 | |
| VinVLInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 84 | |
| OSCAR+ w/ VinVLModel Scale=large2021.01 | 83.98 | |
| VinVLModel Size=Large2021.08 | 83.98 | |
| VinVL-Large# Pretrain Images=5.7M2021.11 | 83.98 | |
| VinVL_largePre-training Images Scale=5.6M2021.11 | 83.98 | |
| VinVL2022.08 | 83.98 | |
| VinVL_L#Images=5.6M, Para.=347M2022.12 | 83.98 | |
| MAPPre-training dataset=<10M images, Model size=Base size2022.10 | 83.48 | |
| METER-Swin-BPretrain Images=4M2022.06 | 83.47 | |
| VLMO-Base# Pretrain Images=4M, Model Size=Base2021.11 | 83.34 | |
| VLMo-BPretrain Images=4M2022.06 | 83.34 | |
| VLMoPre-training image scale=<10M images2022.11 | 83.34 | |
| VLMo-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 83.34 | |
| ManagerTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 83.34 | |
| VLMo-BasePre-training dataset=<10M images, Model size=Base size2022.10 | 83.34 | |
| VLMo_base# Params=175M, Pre-training Data=4M2022.11 | 83.3 | |
| CCLM_basePre-training Data=4M2022.06 | 83.22 | |
| PTP-BLIP#Images=14M, Para.=220M2022.12 | 83.17 | |
| ALBEFPre-training Images=14M2021.07 | 83.14 | |
| ALBEFPre-train #Images=14M2022.01 | 83.14 | |
| ALBEFPre-training Images Scale=14M2021.11 | 83.14 | |
| ALBEF2022.08 | 83.14 | |
| ALBEF(14M)Pre-training image scale=>10M images2022.11 | 83.14 | |
| ALBEF-BASE (14M)# Pre-train Images=14M, Model scale=Base2023.05 | 83.14 | |
| ALBEF2023.05 | 83.14 | |
| ALBEF#Images=14M, Para.=210M2022.12 | 83.14 | |
| ALBEF2022.05 | 83.14 | |
| SimVLM-BasePre-training dataset=>10M images, Model size=Base size2022.10 | 83.14 | |
| ALBEFPre-train # Images=14M2023.12 | 83.14 | |
| ALBEFInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 83.1 | |
| VinVL#Img=6M, Encoder Type=Object Detector (O.D.)2023.04 | 83.1 | |
| VinVLSupervision Level=Supervised, Pre-training Data=Large-Scale Paired Image-Text Data2023.05 | 83.1 | |
| METER# Params=341M, Pre-training Data=4M2022.11 | 83.1 | |
| BLIP_base# Params=240M, Pre-training Data=More Data2022.11 | 83.1 | |
| BridgeTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 83.09 | |
| OSCAR+ w/ VinVLModel Scale=base2021.01 | 83.08 | |
| VinVL-BaseVisual Embed=Region, Time (ms)=~650, Additional pre-training data=GQA, VQAv2, VG-QA, Open Images2021.02 | 83.08 | |
| BLIPPre-train #Images=129M2022.01 | 83.08 | |
| VinVL#Images=6M2022.02 | 83.08 | |
| BLIP-BASE# Pre-train Images=129M, Model scale=Base2023.05 | 83.08 | |
| VinVL_baseModel Scale=base2022.06 | 83.08 | |
| VinVL-BasePre-training dataset=<10M images, Model size=Base size2022.10 | 83.08 | |
| METER-CLIP2021.11 | 83.05 | |
| METERPre-training image scale=<10M images2022.11 | 83.05 | |
| METER-CLIP-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 83.05 | |
| METER2022.05 | 83.05 | |
| METERPre-training dataset=<10M images, Model size=Base size2022.10 | 83.05 | |
| METER-CLIP-VITBASESupervision Level=Supervised, Pre-training Data=Large-Scale Paired Image-Text Data2023.05 | 83 | |
| VL-BEiT# Params=175M, Pre-training Data=4M2022.11 | 82.7 | |
| CoCa-LargeParams (million)=672.12023.10 | 82.6 | |
| METER-Swin2021.11 | 82.47 | |
| METER-Swin-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 82.47 | |
| BLIPPre-train #Images=14M2022.01 | 82.3 | |
| BLIP#Images=14M, Para.=220M2022.12 | 82.3 | |
| BLIP2022.05 | 82.3 | |
| BLIPPre-train # Images=14M2023.12 | 82.3 | |
| FALCON#Images=4M2025.05 | 82.28 | |
| BLIPCapFilt-LPre-train #Images=129M2022.01 | 82.24 | |
| BLIP2022.08 | 82.24 | |
| BLIP_CapFilt-LPretrain Images=129M2022.06 | 82.24 | |
| BLIPPre-training image scale=>10M images2022.11 | 82.24 | |
| BLIP2023.05 | 82.24 | |
| BLIPInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 82.2 | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 82.16 |