Natural Language Visual Reasoning on NLVR2 (dev)
91.51AccuracyBEIT-3
Evaluation Results
| Method | Links | |
|---|---|---|
| BEIT-32022.08 | 91.51 | |
| BEIT-32023.05 | 91.5 | |
| BEiT-3# Params=1.9B2023.01 | 91.5 | |
| VLMO-Large++# Pretrain Images=1.0B2021.11 | 88.62 | |
| ONE-PEACE2023.05 | 87.8 | |
| X-FMbase# Params=327M, Training Data=More Data2023.01 | 87.6 | |
| X-FMbase# Params=327M2023.01 | 86.3 | |
| X2-VLMbase# Params=255M, Training Data=More Data2023.01 | 86.2 | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 86.1 | |
| CoCa2022.08 | 86.1 | |
| CoCa2023.05 | 86.1 | |
| CoCa#Images=3B, Para.=2.1B2022.12 | 86.1 | |
| CoCa# Params=2.1B2023.01 | 86.1 | |
| X2-VLMbaseBackbone=ResNet-50, # Params=255M2023.01 | 85.9 | |
| VLMO-Large# Pretrain Images=4M, Model Size=Large2021.11 | 85.64 | |
| VLMo-LPretrain Images=4M2022.06 | 85.64 | |
| VLMo2022.05 | 85.64 | |
| VLMoInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 85.6 | |
| X-VLMData=16M+, EPIC=true2022.11 | 85.2 | |
| METERData=16M+, EPIC=true2022.11 | 85 | |
| X-VLMData=4M+, EPIC=true2022.11 | 84.6 | |
| FIBER-BPretrain Images=4M2022.06 | 84.59 | |
| mPLUG2022.05 | 84.58 | |
| PTP-BLIP#Images=14M, Para.=220M2022.12 | 84.55 | |
| SimVLM_HUGEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 84.53 | |
| SimVLMModel Size=Huge2021.08 | 84.53 | |
| SimVLM-Huge# Pretrain Images=1.8B2021.11 | 84.53 | |
| SimVLM2022.08 | 84.53 | |
| SimVLM-HPretrain Images=1.8B2022.06 | 84.53 | |
| SimVLMInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 84.5 | |
| SimVLM2023.05 | 84.5 | |
| X-VLMPre-training Images Scale=16M2021.11 | 84.41 | |
| X-VLMPretrain Images=4M2022.06 | 84.41 | |
| X-VLM2023.05 | 84.41 | |
| X-VLMData=16M+, EPIC=false2022.11 | 84.3 | |
| SimVLM-BASE# Pre-train Images=1.8B, Model scale=Base2023.05 | 84.2 | |
| X-VLM# Params=216M2023.01 | 84.2 | |
| X-VLMPre-training Images Scale=4M2021.11 | 84.16 | |
| SimVLMModel Size=Large2021.08 | 84.13 | |
| SimVLM-Large# Pretrain Images=1.8B2021.11 | 84.13 | |
| SimVLM-LARGE# Pre-train Images=1.8B, Model scale=Large2023.05 | 84.13 | |
| SimVLMSize=Large2022.05 | 84.13 | |
| ALBEFData=16M+, EPIC=true2022.11 | 84.1 | |
| SCLPre-training image scale=<10M images2022.11 | 83.63 | |
| METERData=4M+, EPIC=true2022.11 | 83.5 | |
| MAPPre-training dataset=<10M images, Model size=Base size2022.10 | 83.3 | |
| X-VLMData=4M+, EPIC=false2022.11 | 83.3 | |
| ManagerTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 82.81 | |
| VLMbase# Params=175M2023.01 | 82.8 | |
| VLMO-Base# Pretrain Images=4M, Model Size=Base2021.11 | 82.77 | |
| VLMo-BPretrain Images=4M2022.06 | 82.77 | |
| VLMoPre-training image scale=<10M images2022.11 | 82.77 | |
| VLMo-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 82.77 | |
| VLMo-BasePre-training dataset=<10M images, Model size=Base size2022.10 | 82.77 | |
| VinVLInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 82.7 | |
| METERData=16M+, EPIC=false2022.11 | 82.7 | |
| OSCAR+ w/ VinVLModel Scale=large2021.01 | 82.67 | |
| BLIPPre-train #Images=14M2022.01 | 82.67 | |
| VinVL_LARGEPre-training scale=<10M2021.11 | 82.67 | |
| VinVLModel Size=Large2021.08 | 82.67 | |
| VinVL-Large# Pretrain Images=5.7M2021.11 | 82.67 | |
| VinVL_largePre-training Images Scale=5.6M2021.11 | 82.67 | |
| VinVL2022.08 | 82.67 | |
| VinVL_L#Images=5.6M, Para.=347M2022.12 | 82.67 | |
| BLIP#Images=14M, Para.=220M2022.12 | 82.67 | |
| BLIP2022.05 | 82.67 | |
| BLIPPre-train # Images=14M2023.12 | 82.67 | |
| CCLM_basePre-training Data=4M2022.06 | 82.66 | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 82.66 | |
| FALCON#Images=4M2025.05 | 82.61 | |
| ALBEFInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 82.6 | |
| ALBEFData=16M+, EPIC=false2022.11 | 82.6 | |
| ALBEFPre-training Images=14M2021.07 | 82.55 | |
| ALBEFPre-train #Images=14M2022.01 | 82.55 | |
| ALBEFPre-training volume=14M, Pre-training scale=>10M2021.11 | 82.55 | |
| ALBEFPre-training Images Scale=14M2021.11 | 82.55 | |
| ALBEF2022.08 | 82.55 | |
| ALBEF(14M)Pre-training image scale=>10M images2022.11 | 82.55 | |
| ALBEF-BASE (14M)# Pre-train Images=14M, Model scale=Base2023.05 | 82.55 | |
| ALBEF2023.05 | 82.55 | |
| ALBEF#Images=14M, Para.=210M2022.12 | 82.55 | |
| ALBEF2022.05 | 82.55 | |
| SimVLM-BasePre-training dataset=>10M images, Model size=Base size2022.10 | 82.55 | |
| ALBEFPre-train # Images=14M2023.12 | 82.55 | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 82.52 | |
| MAFA#Images=4M2025.05 | 82.52 | |
| BLIPbase# Params=240M2023.01 | 82.5 | |
| BLIPPre-train #Images=129M2022.01 | 82.48 | |
| BLIP-BASE# Pre-train Images=129M, Model scale=Base2023.05 | 82.48 | |
| METERBackbone=CLIP-ViT-BASE, Pre-training scale=<10M2021.11 | 82.33 | |
| METER-CLIP2021.11 | 82.33 | |
| METERPre-training image scale=<10M images2022.11 | 82.33 | |
| METER-CLIP-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 82.33 | |
| METER2022.05 | 82.33 | |
| METERPre-training dataset=<10M images, Model size=Base size2022.10 | 82.33 | |
| METER# Params=341M2023.01 | 82.3 | |
| METERBackbone=Swin-BASE, Pre-training scale=<10M2021.11 | 82.23 | |
| METER-Swin2021.11 | 82.23 | |
| METER-Swin-BPretrain Images=4M2022.06 | 82.23 | |
| METER-Swin-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 82.23 |