Image-to-Text Retrieval on Flickr30k (test)
96.6R@1BLIP
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BLIPPre-train # Images=14M2023.12 | 96.6 | — | — | — | — | — | — | — | — | — | |
| BLIP2026.03 | 96.6 | 99.8 | — | — | — | — | — | — | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 96.2 | — | — | — | — | — | — | — | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 96.1 | — | — | — | — | — | — | — | — | — | |
| FIBER + EQSIMSetting=Fine-tuning with EQSIM regularization2023.03 | 96 | 99.6 | 99.9 | — | — | — | — | — | — | — | |
| ALBEFPre-train # Images=14M2023.12 | 95.9 | — | — | — | — | — | — | — | — | — | |
| ALBEF2026.03 | 95.9 | 99.8 | — | — | — | — | — | — | — | — | |
| GRIT-VLPPre-train # Images=4M, Pre-train Dataset=4M-Noisy, Reproduced=true2023.12 | 95.5 | — | — | — | — | — | — | — | — | — | |
| METER + EQSIMSetting=Fine-tuning with EQSIM regularization2023.03 | 95.3 | 99.6 | 99.9 | — | — | — | — | — | — | — | |
| HADA2023.01 | 95.2 | 99.7 | 100 | 294.9 | — | — | — | — | — | — | |
| CDDSBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 95.2 | 99.9 | — | — | — | — | — | — | — | — | |
| TCLPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 94.9 | — | — | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 94.6 | 99.9 | — | — | — | — | — | — | — | — | |
| BLIP2023.01 | 94.3 | 99.5 | 99.9 | 293.7 | — | — | — | — | — | — | |
| METERSetting=Standard fine-tuning (FT) on Flickr30K2023.03 | 94.3 | 99.6 | 99.9 | — | — | — | — | — | — | — | |
| ALBEFPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 94.3 | — | — | — | — | — | — | — | — | — | |
| BLIPPre-train # Images=4M, Pre-train Dataset=4M-Clean, Reproduced=true2023.12 | 94.3 | — | — | — | — | — | — | — | — | — | |
| UniAdapter2025.03 | 94.2 | 99.5 | 99.7 | — | 1.09 | 44.96 | — | — | — | — | |
| FPET-UniAdapter2025.03 | 94.1 | 99.4 | 99.9 | — | 0.78 | 34.79 | — | — | — | — | |
| VSE++Backbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 94 | 99.5 | — | — | — | — | — | — | — | — | |
| CDDSBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 93.5 | 99.6 | — | — | — | — | — | — | — | — | |
| HADA2023.01 | 93.3 | 99.6 | 100 | 292.9 | — | — | — | — | — | — | |
| FIBERSetting=Standard fine-tuning (FT) on Flickr30K2023.03 | 92.9 | 99.5 | 99.9 | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 92.9 | 99.3 | — | — | — | — | — | — | — | — | |
| ALBEF2023.01 | 92.6 | 99.3 | 99.9 | 291.8 | — | — | — | — | — | — | |
| VSE++Backbone=ViT-224 + BERT, Patch size=14×142026.03 | 92.2 | 99.1 | — | — | — | — | — | — | — | — | |
| FIBERSetting=Direct evaluation after pre-training2023.03 | 91.6 | 99.5 | 99.8 | — | — | — | — | — | — | — | |
| B22023.01 | 91.4 | 99.5 | 99.7 | 290.6 | — | — | — | — | — | — | |
| METERSetting=Direct evaluation after pre-training2023.03 | 90.9 | 98.3 | 99.5 | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-L/14, Zero-shot=true, Fine-tuned with=DOCCI2025.03 | 90.8 | 98.8 | — | — | — | — | — | 99.9 | 100 | — | |
| B12023.01 | 90.7 | 99 | 99.6 | 289.3 | — | — | — | — | — | — | |
| DivE (Ours)Backbone=ResNeXt-101 + BERT, CA=false, Ensemble=true2022.11 | 90.6 | 99 | 99.6 | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 90.5 | 98.9 | 99.7 | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 90.3 | 98.5 | 99.8 | — | — | — | — | — | — | — | |
| Long-CLIPBackbone=ViT-L/14, Zero-shot=true2025.03 | 90 | 98.9 | — | — | — | — | — | 99.9 | 100 | — | |
| SCANBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 90 | 98.5 | — | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 90 | 98.2 | 99.5 | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 90 | 99.3 | 99.9 | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 89.6 | 99.2 | 99.8 | — | — | — | — | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 89.1 | 98.5 | 99.6 | — | — | — | — | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 89.1 | 98.4 | 99.5 | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-L/14, Zero-shot=true, Fine-tuned with=DCI2025.03 | 89.1 | 98.3 | — | — | — | — | — | 100 | 100 | — | |
| SigLIPZero-shot=true2025.09 | 88.9 | — | — | — | — | — | — | — | — | — | |
| DivE (Ours)Backbone=ResNeXt-101 + BERT, CA=false2022.11 | 88.8 | 98.5 | 99.6 | — | — | — | — | — | — | — | |
| VSE∞Backbone=ResNeXt-101 + BERT, CA=false, Ensemble=true2022.11 | 88.7 | 98.9 | 99.8 | — | — | — | — | — | — | — | |
| CLIP-L/14 + RECOBackbone=CLIP-L/14, # par.=435M, Zero-shot=true2023.06 | 88.5 | — | — | — | — | — | — | — | — | — | |
| VSE∞Backbone=ResNeXt-101 + BERT, CA=false2022.11 | 88.4 | 98.3 | 99.5 | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 88.4 | 98.5 | 99.6 | — | — | — | — | — | — | — | |
| SCANBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 88.2 | 98.1 | — | — | — | — | — | — | — | — | |
| CLIPFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 88.1 | 98.2 | 99.6 | — | — | — | — | — | — | — | |
| CLIP2023.01 | 88 | 98.7 | 99.4 | 286.1 | — | — | — | — | — | — | |
| VILLA2023.01 | 87.9 | 97.2 | 98.8 | 283.9 | — | — | — | — | — | — | |
| FFF-ViT-B/32Pre-train dataset=Open30M, Zero-shot=true2024.05 | 87.9 | 99.2 | 99.6 | — | — | — | — | — | — | — | |
| VILLAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 87.9 | — | — | — | — | — | — | — | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 87.8 | 99.1 | 99.8 | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 87.8 | 97.5 | 98.8 | — | — | — | — | — | — | — | |
| LAION-CLIPBackbone=ViT-L, Evaluation Protocol=zero-shot2024.09 | 87.6 | — | — | — | — | — | — | — | — | — | |
| LAION-CLIP ViT-LBackbone=ViT-L, Training Samples=400M2024.09 | 87.6 | — | — | — | — | — | — | — | — | — | |
| CLIP-L/14Backbone=CLIP-L/14, # par.=428M, Zero-shot=true2023.06 | 87.5 | — | — | — | — | — | — | — | — | — | |
| FFF-ViT-B/32Pre-train dataset=Open70M, Zero-shot=true2024.05 | 87.5 | 98.1 | 99.3 | — | — | — | — | — | — | — | |
| DINOv2-ARLEvaluation Protocol=zero-shot2024.09 | 87.5 | — | — | — | — | — | — | — | — | — | |
| DINOv2-ARLTraining Samples=20M2024.09 | 87.5 | — | — | — | — | — | — | — | — | — | |
| CLIPvisual encoder=ViT-L/14-336, visual FLOPS=191.0G2023.12 | 87.4 | — | — | — | — | — | — | — | — | — | |
| UNITER2023.01 | 87.3 | 98 | 99.2 | 284.5 | — | — | — | — | — | — | |
| 3TModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 87.3 | — | — | — | — | — | — | — | — | — | |
| UNITERPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 87.3 | — | — | — | — | — | — | — | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 87.3 | 97.9 | 99.1 | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 86.8 | 98 | 99.3 | — | — | — | — | — | — | — | |
| CLIPFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 86.8 | 98.3 | 99.8 | — | — | — | — | — | — | — | |
| Align# par.=247M, Zero-shot=true2023.06 | 86.7 | — | — | — | — | — | — | — | — | — | |
| SOHOBackbone=R1012021.04 | 86.5 | 98.1 | 99.3 | — | — | — | — | — | — | — | |
| SOHO2026.03 | 86.5 | 98.1 | — | — | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/14, Zero-shot=true2025.03 | 86.4 | 97.5 | — | — | — | — | — | 99.9 | 100 | — | |
| PyramidCLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true2022.04 | 86.3 | 98 | — | — | — | — | — | — | — | — | |
| MaxMatchBackbone=Faster R-CNN + BERT, Cross-Attention=false, Ensemble=true2025.06 | 86.2 | 95.7 | 98.4 | 527.1 | — | — | — | — | — | — | |
| Unicoder-VL2021.04 | 86.2 | 96.3 | 99 | — | — | — | — | — | — | — | |
| LongCLIPZero-shot=true2025.09 | 86.2 | — | — | — | — | — | — | — | — | — | |
| UNITERBackbone=R1012021.04 | 85.9 | 97.1 | 98.8 | — | — | — | — | — | — | — | |
| Long-CLIPBackbone=ViT-B/16, Zero-shot=true2025.03 | 85.9 | 98.5 | — | — | — | — | — | 99.9 | 100 | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true2022.04 | 85.6 | 97.7 | — | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 85.5 | 98.4 | 99.1 | — | — | — | — | — | — | — | |
| FFF-ViT-B/32Pre-train dataset=YFCC-15M, Zero-shot=true2024.05 | 85.3 | 97.5 | 99.4 | — | — | — | — | — | — | — | |
| OpenAI-CLIPBackbone=ViT-L, Evaluation Protocol=zero-shot2024.09 | 85.2 | — | — | — | — | — | — | — | — | — | |
| OpenAI-CLIP ViT-LBackbone=ViT-L, Training Samples=400M2024.09 | 85.2 | — | — | — | — | — | — | — | — | — | |
| CLIPFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 85.2 | 97.3 | 99.1 | — | — | — | — | — | — | — | |
| CLIPvisual encoder=ViT-L/14-224, visual FLOPS=81.1G2023.12 | 85.1 | — | — | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-B/16, Zero-shot=true, Fine-tuned with=DOCCI2025.03 | 85.1 | 96.7 | — | — | — | — | — | 99.6 | 99.9 | — | |
| NegCLIPZero-shot=true2025.09 | 85.1 | — | — | — | — | — | — | — | — | — | |
| Baseline (CLIP-style)Model scale=g scale, Training dataset=Text-Filtered WebLI2023.05 | 85 | — | — | — | — | — | — | — | — | — | |
| CLIPBackbone=ViT-224-Large + BERT-Large, Patch size=16×16, Zero-shot=true2026.03 | 85 | 97.7 | — | — | — | — | — | — | — | — | |
| TinyCLIPvisual encoder=ViT-63M/32-224, visual FLOPS=2.0G2023.12 | 84.9 | — | — | — | — | — | — | — | — | — | |
| DreamLIPZero-shot=true2025.09 | 84.7 | — | — | — | — | — | — | — | — | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 84.6 | 97.3 | — | — | — | — | — | — | — | — | |
| SF-LCLIPvisual encoder=SF-LCLIP, tokens=144, visual FLOPS=33.9G2023.12 | 84.6 | — | — | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-B/16, Zero-shot=true, Fine-tuned with=DCI2025.03 | 84.6 | 96.8 | — | — | — | — | — | 99.8 | 100 | — | |
| DINOv2-MpNetEvaluation Protocol=zero-shot2024.09 | 84.6 | — | — | — | — | — | — | — | — | — | |
| DINOv2-MpNetBackbone=MpNet, Training Samples=20M2024.09 | 84.6 | — | — | — | — | — | — | — | — | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 84.5 | 97.4 | — | — | — | — | — | — | — | — | |
| MaxMatchBackbone=Faster R-CNN + BERT, Cross-Attention=false2025.06 | 84.2 | 96.1 | 97.9 | 520.8 | — | — | — | — | — | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset=143M, Zero-shot=true2022.04 | 84.2 | 96.4 | — | — | — | — | — | — | — | — |