Visual Entailment on SNLI-VE (test)
91.2Overall AccuracyOFA
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OFA2022.02 | 91.2 | — | — | — | — | — | — | — | — | — | — | |
| OFALarge2022.02 | 90.2 | — | — | — | — | — | — | — | — | — | — | |
| OFAInput Modality=image and text premises, Evaluation Protocol=lightweight finetuning2022.05 | 90.2 | — | — | — | — | — | — | — | — | — | — | |
| OFA2022.05 | 90.2 | — | — | — | — | — | — | — | — | — | — | |
| mPLUG2022.05 | 89.29 | — | — | — | — | — | — | — | — | — | — | |
| OFABase2022.02 | 89.2 | — | — | — | — | — | — | — | — | — | — | |
| OFA BaseEncoder Layers=6, Decoder Layers=62022.11 | 89.2 | — | — | — | 1 | — | — | — | — | — | — | |
| MuEBackbone=OFA Base2022.11 | 88.5 | — | — | — | -50 | — | — | — | — | — | — | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 87.1 | — | — | — | — | — | — | — | — | — | — | |
| OFAMedium2022.02 | 87 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM_HUGEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 86.32 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMModel Size=Huge2021.08 | 86.32 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMModel Size=Huge, Pre-trained Data=1.8B image-text pairs, Backbone=ViT-Huge2022.02 | 86.3 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 86.3 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMModel Size=Large2021.08 | 85.62 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM-LARGE# Pre-train Images=1.8B, Model scale=Large2023.05 | 85.62 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMSize=Large2022.05 | 85.62 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-LargeParams (million)=672.12023.10 | 85.6 | — | — | — | — | — | — | — | — | — | — | |
| OFATiny2022.02 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| OFA TinyEncoder Layers=4, Decoder Layers=42022.11 | 85.2 | — | — | — | -33 | — | — | — | — | — | — | |
| PABEEBackbone=OFA2022.11 | 85.2 | — | — | — | -15.3 | — | — | — | — | — | — | |
| SOHOBackbone=R1012021.04 | 84.95 | — | — | — | — | — | — | — | — | — | — | |
| SOHOModel Size=Base2021.08 | 84.95 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-Tiny_12 (OPTIMA)Teacher=CoCa-Large, Params (million)=101.8, Inference Speedup=3.4x2023.10 | 84.3 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM_BASEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 84.15 | — | — | — | — | — | — | — | — | — | — | |
| SimVLMModel Size=Base2021.08 | 84.15 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF (14M)Pre-training dataset=>10M images, Model size=Base size2022.10 | 84.15 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-Tiny_12Teacher=CoCa-Large, Params (million)=101.8, Inference Speedup=3.4x2023.10 | 83.9 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-BaseParams (million)=293.12023.10 | 83.6 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-Tiny_6 (OPTIMA)Teacher=CoCa-Large, Params (million)=54.5, Inference Speedup=6.2x2023.10 | 82.3 | — | — | — | — | — | — | — | — | — | — | |
| CoCa-Tiny_6Teacher=CoCa-Large, Params (million)=54.5, Inference Speedup=6.2x2023.10 | 82 | — | — | — | — | — | — | — | — | — | — | |
| ManagerTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 81.44 | — | — | — | — | — | — | — | — | — | — | |
| MAPPre-training dataset=<10M images, Model size=Base size2022.10 | 81.39 | — | — | — | — | — | — | — | — | — | — | |
| METER2022.02 | 81.2 | — | — | — | — | — | — | — | — | — | — | |
| METER-CLIP-VITBASESupervision Level=Supervised, Pre-training Data=Large-Scale Paired Image-Text Data2023.05 | 81.2 | — | — | — | — | — | — | — | — | — | — | |
| METERBackbone=CLIP-ViT-BASE, Pre-training scale=<10M2021.11 | 81.19 | — | — | — | — | — | — | — | — | — | — | |
| METER-CLIP-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 81.19 | — | — | — | — | — | — | — | — | — | — | |
| BridgeTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 81.19 | — | — | — | — | — | — | — | — | — | — | |
| METER2022.05 | 81.19 | — | — | — | — | — | — | — | — | — | — | |
| METERPre-training dataset=<10M images, Model size=Base size2022.10 | 81.19 | — | — | — | — | — | — | — | — | — | — | |
| METERFine-tuning protocol=original fine-tuned2023.05 | 81.1 | — | — | — | — | — | — | — | — | — | — | |
| Knowledge-CLIPmode=Fine-tuning2022.10 | 80.97 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-training Images=14M2021.07 | 80.91 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-training volume=14M, Pre-training scale=>10M2021.11 | 80.91 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF-BASE (14M)# Pre-train Images=14M, Model scale=Base2023.05 | 80.91 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF2022.05 | 80.91 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM-BasePre-training dataset=>10M images, Model size=Base size2022.10 | 80.91 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF2022.02 | 80.9 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.9 | — | — | — | — | — | — | — | — | — | — | |
| UNIMO_LARGEPre-training scale=<10M2021.11 | 80.63 | — | — | — | — | — | — | — | — | — | — | |
| UNIMO-LARGE# Pre-train Images=4M, Model scale=Large2023.05 | 80.63 | — | — | — | — | — | — | — | — | — | — | |
| UNIMO2022.05 | 80.63 | — | — | — | — | — | — | — | — | — | — | |
| UNIMOModel Size=large2020.12 | 80.63 | — | — | — | — | — | — | — | — | — | — | |
| UNIMO2022.02 | 80.6 | — | — | — | — | — | — | — | — | — | — | |
| SoftMask++#Img=4M, Encoder Type=Detector-Free (D.F.)2023.04 | 80.6 | — | — | — | — | — | — | — | — | — | — | |
| METERBackbone=Swin-BASE, Pre-training scale=<10M2021.11 | 80.45 | — | — | — | — | — | — | — | — | — | — | |
| METER-Swin-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80.45 | — | — | — | — | — | — | — | — | — | — | |
| CODIS#Img=4M, Encoder Type=Detector-Free (D.F.)2023.04 | 80.4 | — | — | — | — | — | — | — | — | — | — | |
| MADBase Model=VILLA, Teacher Model (VE)=CLIP-V, Teacher Model (TE)=CLIP-T, Training Samples=Full2022.04 | 80.32 | — | — | — | — | — | — | — | — | — | — | |
| MADBase Model=VILLA, Teacher Model (TE)=RoBERTa, Training Samples=Full2022.04 | 80.31 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-training Images=4M2021.07 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-training volume=4M, Pre-training scale=<10M2021.11 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF#Images=4M2022.02 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFmode=Fine-tuning2022.10 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF#Img=4M, Encoder Type=Detector-Free (D.F.)2023.04 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| TCL#Img=4M, Encoder Type=Detector-Free (D.F.)2023.04 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| METER + PuMerFine-tuning protocol=PuMer fine-tuned2023.05 | 80.3 | — | — | — | — | 2.07 | 43 | — | — | — | — | |
| ALBEFSupervision Level=Supervised, Pre-training Data=Large-Scale Paired Image-Text Data2023.05 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF (4M)Pre-training dataset=<10M images, Model size=Base size2022.10 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFParams (million)=241.02023.10 | 80.3 | — | — | — | — | — | — | — | — | — | — | |
| TCL#Images=4M2022.02 | 80.29 | — | — | — | — | — | — | — | — | — | — | |
| MADBase Model=UNITER, Teacher Model (VE)=CLIP-V, Teacher Model (TE)=CLIP-T, Training Samples=Full2022.04 | 80.23 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-ViLPre-training scale=<10M, Backbone=ResNet50x42021.11 | 80.2 | — | — | — | — | — | — | — | — | — | — | |
| BaselineBase Model=CLIP-ViLp, Teacher Model (VE)=CLIP-V, Training Samples=Full2022.04 | 80.2 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-ViLInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.2 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-ViL2022.05 | 80.2 | — | — | — | — | — | — | — | — | — | — | |
| MDBase Model=UNITER, Teacher Model (VE)=CLIP-V, Teacher Model (TE)=CLIP-T, Training Samples=Full2022.04 | 80.16 | — | — | — | — | — | — | — | — | — | — | |
| Geometric Consistency LossVision Encoder=ViT-B/16, Text Encoder=BERT-base2023.03 | 80.16 | — | — | — | — | — | — | — | — | — | — | |
| CODISVision Encoder=ViT-B/16, Text Encoder=BERT-base2023.03 | 80.13 | — | — | — | — | — | — | — | — | — | — | |
| VILLAModel size=Large2020.06 | 80.02 | — | — | — | — | — | — | — | — | — | — | |
| VILLA_LARGEPre-training scale=<10M2021.11 | 80.02 | — | — | — | — | — | — | — | — | — | — | |
| VillaModel Size=Large2021.08 | 80.02 | — | — | — | — | — | — | — | — | — | — | |
| VillaModel Size=large2020.12 | 80.02 | — | — | — | — | — | — | — | — | — | — | |
| CLIPmode=Fine-tuning2022.10 | 80.01 | — | — | — | — | — | — | — | — | — | — | |
| VILLA2022.02 | 80 | — | — | — | — | — | — | — | — | — | — | |
| Brownian Bridge LossVision Encoder=ViT-B/16, Text Encoder=BERT-base2023.03 | 79.95 | — | — | — | — | — | — | — | — | — | — | |
| ALBEFVision Encoder=ViT-B/16, Text Encoder=BERT-base2023.03 | 79.91 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF BaseParams=581, TFLOPS=7.052023.10 | 79.79 | — | — | — | — | — | — | — | — | — | — | |
| Deep Feature Separation LossVision Encoder=ViT-B/16, Text Encoder=BERT-base2023.03 | 79.61 | — | — | — | — | — | — | — | — | — | — | |
| ALBEF LORAParams=644, TFLOPS=7.142023.10 | 79.53 | — | — | — | — | — | — | — | — | — | — | |
| UNITER2022.02 | 79.4 | — | — | — | — | — | — | — | — | — | — | |
| UNITERInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 79.4 | — | — | — | — | — | — | — | — | — | — | |
| UNITERModel size=Large2020.06 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| UNITER_LARGEPre-training scale=<10M2021.11 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| UNITERModel Size=Large2021.08 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| UNITER-LARGE# Pre-train Images=4M, Model scale=Large2023.05 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| UNITER2022.05 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| UNITERModel Size=large2020.12 | 79.38 | — | — | — | — | — | — | — | — | — | — | |
| BaselineBase Model=VILLA, Training Samples=Full2022.04 | 79.32 | — | — | — | — | — | — | — | — | — | — |