Visual Entailment on SNLI-VE (dev)
91AccuracyOFA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OFA2022.02 | 91 | — | |
| OFALarge2022.02 | 90.3 | — | |
| OFAInput Modality=image and text premises, Evaluation Protocol=lightweight finetuning2022.05 | 90.3 | — | |
| OFA2022.05 | 90.3 | — | |
| mPLUG2022.05 | 89.45 | — | |
| OFABase2022.02 | 89.3 | — | |
| OFA BaseEncoder Layers=6, Decoder Layers=62022.11 | 89.3 | 1 | |
| MuEBackbone=OFA Base2022.11 | 88.7 | -50 | |
| CoCaInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 87 | — | |
| OFAMedium2022.02 | 86.6 | — | |
| SimVLM_HUGEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 86.21 | — | |
| SimVLMModel Size=Huge2021.08 | 86.21 | — | |
| SimVLMModel Size=Huge, Pre-trained Data=1.8B image-text pairs, Backbone=ViT-Huge2022.02 | 86.2 | — | |
| SimVLMInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 86.2 | — | |
| SimVLMModel Size=Large2021.08 | 85.68 | — | |
| SimVLM-LARGE# Pre-train Images=1.8B, Model scale=Large2023.05 | 85.68 | — | |
| SimVLMSize=Large2022.05 | 85.68 | — | |
| OFATiny2022.02 | 85.3 | — | |
| OFA TinyEncoder Layers=4, Decoder Layers=42022.11 | 85.3 | -33 | |
| PABEEBackbone=OFA2022.11 | 85.3 | -15.3 | |
| SOHOModel Size=Base2021.08 | 85 | — | |
| SimVLM_BASEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 84.2 | — | |
| SimVLMModel Size=Base2021.08 | 84.2 | — | |
| METERData=16M+, EPIC=true2022.11 | 82.1 | — | |
| METERData=16M+, EPIC=false2022.11 | 81.7 | — | |
| METERData=4M+, EPIC=true2022.11 | 81.6 | — | |
| METERData=4M+, EPIC=false2022.11 | 81.4 | — | |
| ALBEFData=16M+, EPIC=true2022.11 | 81.3 | — | |
| ManagerTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 81.26 | — | |
| UNIMO_LARGEPre-training scale=<10M2021.11 | 81.11 | — | |
| BridgeTower-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 81.11 | — | |
| UNIMO-LARGE# Pre-train Images=4M, Model scale=Large2023.05 | 81.11 | — | |
| UNIMO2022.05 | 81.11 | — | |
| UNIMO2022.02 | 81.1 | — | |
| METER2022.02 | 80.9 | — | |
| METERBackbone=CLIP-ViT-BASE, Pre-training scale=<10M2021.11 | 80.86 | — | |
| METER-CLIP-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80.86 | — | |
| METER2022.05 | 80.86 | — | |
| ALBEFPre-training volume=14M, Pre-training scale=>10M2021.11 | 80.8 | — | |
| ALBEF2022.02 | 80.8 | — | |
| ALBEFInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.8 | — | |
| ALBEF-BASE (14M)# Pre-train Images=14M, Model scale=Base2023.05 | 80.8 | — | |
| ALBEF2022.05 | 80.8 | — | |
| ALBEFData=16M+, EPIC=false2022.11 | 80.8 | — | |
| CLIP-ViLpVisual Encoder=CLIP-Res50x4, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 80.61 | — | |
| CLIP-ViLPre-training scale=<10M, Backbone=ResNet50x42021.11 | 80.61 | — | |
| METERBackbone=Swin-BASE, Pre-training scale=<10M2021.11 | 80.61 | — | |
| METER-Swin-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80.61 | — | |
| CLIP-ViL2022.05 | 80.61 | — | |
| CLIP-ViLInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 80.6 | — | |
| ALBEFData=4M+, EPIC=true2022.11 | 80.6 | — | |
| VILLA2022.02 | 80.2 | — | |
| VILLA_LARGEPre-training scale=<10M2021.11 | 80.18 | — | |
| VillaModel Size=Large2021.08 | 80.18 | — | |
| ALBEFPre-training volume=4M, Pre-training scale=<10M2021.11 | 80.14 | — | |
| ALBEF-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80.14 | — | |
| ALBEFData=4M+, EPIC=false2022.11 | 80.1 | — | |
| UNIMO-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 80 | — | |
| UNITER2022.02 | 79.4 | — | |
| UNITERInput Modality=image only, Evaluation Protocol=lightweight finetuning2022.05 | 79.4 | — | |
| UNITER_LARGEPre-training scale=<10M2021.11 | 79.39 | — | |
| UNITERModel Size=Large2021.08 | 79.39 | — | |
| UNITER-LARGE# Pre-train Images=4M, Model scale=Large2023.05 | 79.39 | — | |
| UNITER2022.05 | 79.39 | — | |
| DeeBERTBackbone=OFA2022.11 | 78.9 | -15 | |
| CLIP-ViLpVisual Encoder=CLIP-Res50, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 78.64 | — | |
| UNITERVisual Encoder=BUTD-Res101, V&L Pretrain Data=6.5M2021.07 | 78.59 | — | |
| UNITER-BASE# Pre-train Images=4M, Model scale=Base2023.05 | 78.59 | — | |
| EfficientParams=39.2M, Training=20 Epochs2026.04 | 75.1 | — | |
| VL-T5Model Size=Base2021.08 | 73.6 | — | |
| OFABasezero-shot=true2022.02 | 49.71 | — |