Image-to-text retrieval on Flickr30K 1K Karpathy (test)
97.2R@1Florence-Huge
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Florence-HugePretrain Images=900M, Model Size Category=Huge-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 97.2 | 99.9 | — | |
| VLMO-Large++Pretrain Images=1.0B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product), large batch size=true2021.11 | 96.8 | 100 | 100 | |
| VLMO-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 95.3 | 99.9 | 100 | |
| ALIGN-LargePretrain Images=1.8B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 95.3 | 99.8 | 100 | |
| BEIT-3Evaluation Mode=Zero-shot2022.08 | 94.9 | 99.9 | 100 | |
| ALBEF-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion reranking2021.11 | 94.3 | 99.4 | 99.8 | |
| CoCaEvaluation Mode=Zero-shot2022.08 | 92.5 | 99.5 | 99.9 | |
| VLMO-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 92.3 | 99.4 | 99.9 | |
| COTS# PT Pairs=15.3M, Architecture=Two-Stream, Visual Encoder Pre-training=standard, Ensemble=true2022.04 | 91.7 | 99 | 99.9 | |
| FlorenceEvaluation Mode=Zero-shot2022.08 | 90.9 | 99.1 | — | |
| COTS# PT Pairs=15.3M, Architecture=Two-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 90.6 | 98.7 | 99.7 | |
| FILIPEvaluation Mode=Zero-shot2022.08 | 89.8 | 99.2 | 99.8 | |
| FlamingoEvaluation Mode=Zero-shot2022.08 | 89.3 | 98.8 | 99.7 | |
| COOKIE# PT Pairs=5.9M, Architecture=Two-Stream, Visual Encoder Pre-training=940M tagged images, Ensemble=true2022.04 | 89 | 98.9 | 99.7 | |
| VSE∞# PT Pairs=N/A, Architecture=Two-Stream, Visual Encoder Pre-training=940M tagged images, Ensemble=true2022.04 | 88.7 | 98.9 | 99.8 | |
| ALIGNEvaluation Mode=Zero-shot2022.08 | 88.6 | 98.7 | 99.7 | |
| COTS# PT Pairs=5.3M, Architecture=Two-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 88.2 | 98.5 | 99.7 | |
| VLMO-Large++Pretrain Images=1.0B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product), large batch size=true2021.11 | 88.1 | 98.4 | 99.3 | |
| CLIPEvaluation Mode=Zero-shot2022.08 | 88 | 98.7 | 99.4 | |
| VILLA-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 87.9 | 97.5 | 98.8 | |
| Florence-HugePretrain Images=900M, Model Size Category=Huge-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 87.9 | 98.1 | — | |
| UNITER-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 87.3 | 98 | 99.2 | |
| Pixel-BERT-X152# PT Pairs=5.6M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 87 | 98.9 | 99.5 | |
| ERNIE-ViL-base# PT Pairs=3.8M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 86.7 | 97.8 | 99 | |
| VILLA-Base# PT Pairs=9.6M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 86.6 | 97.9 | 99.2 | |
| VILLA-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 86.6 | 97.9 | 99.2 | |
| Unicoder-VL# PT Pairs=3.8M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 86.2 | 96.3 | 99 | |
| UNITER-Base# PT Pairs=9.6M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 85.9 | 97.1 | 98.8 | |
| UNITER-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 85.9 | 97.1 | 98.8 | |
| VinVL-LargePretrain Images=5.7M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 84.9 | 97.4 | 98.6 | |
| ALIGN-LargePretrain Images=1.8B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 84.9 | 97.4 | 98.6 | |
| COOKIE# PT Pairs=5.9M, Architecture=Two-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 84.7 | 96.9 | 98.3 | |
| VLMO-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 84.5 | 97.3 | 98.6 | |
| LightningDOT# PT Pairs=9.5M, Architecture=Two-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 83.9 | 97.2 | 98.6 | |
| ViLT# PT Pairs=9.9M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 83.5 | 96.7 | 98.6 | |
| ViLT-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 83.5 | 96.7 | 98.6 | |
| ALBEF-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion reranking2021.11 | 82.8 | 96.7 | 98.4 | |
| BEIT-3Evaluation Mode=Zero-shot2022.08 | 81.5 | 95.6 | 97.8 | |
| CoCaEvaluation Mode=Zero-shot2022.08 | 80.4 | 95.7 | 97.7 | |
| FlamingoEvaluation Mode=Zero-shot2022.08 | 79.5 | 95.3 | 97.9 | |
| VLMO-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 79.3 | 95.7 | 97.8 | |
| FlorenceEvaluation Mode=Zero-shot2022.08 | 76.7 | 93.6 | — | |
| VILLA-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 76.3 | 94.2 | 96.8 | |
| Pixel-BERT-R50# PT Pairs=5.6M, Architecture=Single-Stream, Visual Encoder Pre-training=standard, Ensemble=false2022.04 | 75.7 | 94.7 | 97.1 | |
| ALIGNEvaluation Mode=Zero-shot2022.08 | 75.7 | 93.8 | 96.8 | |
| UNITER-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 75.6 | 94.1 | 96.8 | |
| FILIPEvaluation Mode=Zero-shot2022.08 | 75 | 93.4 | 96.3 | |
| VILLA-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 74.7 | 92.9 | 95.8 | |
| UNITER-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 72.5 | 92.4 | 96.1 | |
| CLIPEvaluation Mode=Zero-shot2022.08 | 68.7 | 90.6 | 95.2 | |
| FLAVAEvaluation Mode=Zero-shot2022.08 | 67.7 | 94 | — | |
| FLAVAEvaluation Mode=Zero-shot2022.08 | 65.2 | 89.4 | — | |
| ViLT-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion encoder2021.11 | 64.4 | 88.7 | 93.8 | |
| MosaiCLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 44.5 | — | — | |
| NegCLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 38.6 | — | — | |
| CLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 36.2 | — | — | |
| MosaiCLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 29.5 | — | — | |
| CLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 24.1 | — | — | |
| NegCLIPBackbone=Swin-Tiny, Pre-training data=YFCC-15M2023.05 | 23.3 | — | — |