Text-to-Image Retrieval on Flickr30k (test)
89.7Recall@1BLIP-2
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BLIP-2Throughput=1.68/s2024.06 | 89.7 | 98.1 | 98.9 | — | — | — | — | — | — | — | — | — | — | |
| BLIPPre-train # Images=14M2023.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLIP2026.03 | 87.2 | 97.5 | — | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-train # Images=14M2023.12 | 85.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ALBEF2026.03 | 85.6 | 97.5 | — | — | — | — | — | — | — | — | — | — | — | |
| HADA2023.01 | 85.3 | 97.24 | 98.72 | 281.26 | — | — | — | — | 576.16 | 3.94 | — | — | — | |
| CDDSBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 85.3 | 98 | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-GThroughput=2.03/s2024.06 | 85 | 97 | 98.6 | — | — | — | — | — | — | — | — | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 84.9 | 97.3 | — | — | — | — | — | — | — | — | — | — | — | |
| InsertQuantBackbone=CLIP ViT-L, Bits (W-A-KV)=16-16-16, Evaluation Protocol=zero-shot retrieval2026.06 | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InsertQuantBackbone=CLIP ViT-L, Bits (W-A-KV)=8-8s-8, Evaluation Protocol=zero-shot retrieval2026.06 | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaselineBackbone=CLIP ViT-L, Bits (W-A-KV)=16-16-16, Evaluation Protocol=zero-shot retrieval2026.06 | 84.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TCLPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 84 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InsertQuantBackbone=CLIP ViT-L, Bits (W-A-KV)=4-8s-8, Evaluation Protocol=zero-shot retrieval2026.06 | 83.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniAdapter2025.03 | 83.6 | 96.6 | 98.2 | — | — | 1.09 | 44.96 | — | — | — | — | — | — | |
| FIBER + EQSIMSetting=Fine-tuning with EQSIM regularization2023.03 | 83.56 | 96.78 | 98.28 | — | — | — | — | — | — | — | — | — | — | |
| BLIP2023.01 | 83.54 | 96.66 | 98.32 | 278.52 | — | — | — | — | 572.22 | 0 | — | — | — | |
| VSE++Backbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 83.4 | 96.4 | — | — | — | — | — | — | — | — | — | — | — | |
| RTNBackbone=CLIP ViT-L, Bits (W-A-KV)=8-8s-8, Evaluation Protocol=zero-shot retrieval2026.06 | 83.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FPET-UniAdapter2025.03 | 83 | 96 | 98 | — | — | 0.78 | 34.79 | — | — | — | — | — | — | |
| GRIT-VLPPre-train # Images=4M, Pre-train Dataset=4M-Noisy, Reproduced=true2023.12 | 82.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ALBEFPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RTNBackbone=CLIP ViT-L, Bits (W-A-KV)=4-8s-8, Evaluation Protocol=zero-shot retrieval2026.06 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLIPPre-train # Images=4M, Pre-train Dataset=4M-Clean, Reproduced=true2023.12 | 82.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| METERSetting=Standard fine-tuning (FT) on Flickr30K2023.03 | 82.22 | 96.34 | 98.36 | — | — | — | — | — | — | — | — | — | — | |
| METER + EQSIMSetting=Fine-tuning with EQSIM regularization2023.03 | 82.16 | 94.7 | 96.64 | — | — | — | — | — | — | — | — | — | — | |
| CARTThroughput=105.8/s2024.06 | 81.8 | 96.1 | 98.4 | — | — | — | — | — | — | — | — | — | — | |
| CART2024.06 | 81.8 | 96.1 | 98.4 | — | 88 | — | — | — | — | — | — | — | — | |
| FIBERSetting=Standard fine-tuning (FT) on Flickr30K2023.03 | 81.44 | 96.72 | 98.48 | — | — | — | — | — | — | — | — | — | — | |
| InsertQuantBackbone=CLIP ViT-L, Bits (W-A-KV)=4-6s-6, Evaluation Protocol=zero-shot retrieval2026.06 | 81.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HADA2023.01 | 81.36 | 95.94 | 98.02 | 275.32 | — | — | — | — | — | — | — | — | — | |
| CDDSBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 81.2 | 96.4 | — | — | — | — | — | — | — | — | — | — | — | |
| SCANBackbone=ViT-224-Large + BERT-Large, Patch size=16×162026.03 | 81 | 95.9 | — | — | — | — | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 80.6 | 95.5 | — | — | — | — | — | — | — | — | — | — | — | |
| VSE++Backbone=ViT-224 + BERT, Patch size=14×142026.03 | 80.5 | 95.6 | — | — | — | — | — | — | — | — | — | — | — | |
| ALBEF2023.01 | 79.76 | 95.3 | 97.72 | 272.78 | — | — | — | — | — | — | — | — | — | |
| B22023.01 | 79.64 | 95.34 | 97.46 | 272.44 | — | — | — | — | — | — | — | — | — | |
| METERSetting=Direct evaluation after pre-training2023.03 | 79.6 | 94.96 | 97.28 | — | — | — | — | — | — | — | — | — | — | |
| MACPT Dataset=CC3M, WV2M2022.12 | 79.3 | 94.7 | 97.2 | — | — | — | — | — | — | — | — | — | — | |
| FIBERSetting=Direct evaluation after pre-training2023.03 | 79.26 | 95.7 | 97.92 | — | — | — | — | — | — | — | — | — | — | |
| B12023.01 | 79.08 | 94.5 | 96.94 | 270.52 | — | — | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 78.74 | 95.1 | 97.42 | — | — | — | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 78.58 | 95.16 | 97.6 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 77.36 | 94.22 | 97.1 | — | — | — | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 77.3 | 94.38 | 97.14 | — | — | — | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 77.28 | 94.58 | 97.36 | — | — | — | — | — | — | — | — | — | — | |
| VILLAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 76.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VILLA2023.01 | 76.26 | 94.24 | 96.84 | 267.34 | — | — | — | — | — | — | — | — | — | |
| Long-CLIPBackbone=ViT-L/14, Zero-shot=true2025.03 | 76.22 | 93.54 | — | — | — | — | — | — | — | — | 98.36 | 99.28 | — | |
| VSE∞Backbone=ResNeXt-101 + BERT, CA=false, Ensemble=true2022.11 | 76.1 | 94.5 | 97.1 | — | — | — | — | — | — | — | — | — | — | |
| OSCARObject Detector=true, Input Size=Full, Pre-training Data=COCO+CC+SBU+GQA, Zero-shot Setting=false2021.03 | 75.9 | 93.3 | 96.6 | — | — | — | — | — | — | — | — | — | — | |
| DivE (Ours)Backbone=ResNeXt-101 + BERT, CA=false, Ensemble=true2022.11 | 75.9 | 94.7 | 97.3 | — | — | — | — | — | — | — | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 75.9 | 93.1 | 96.8 | — | — | — | — | — | — | — | — | — | — | |
| UNITERObject Detector=true, Input Size=Full, Pre-training Data=COCO+CC+SBU+VG, Zero-shot Setting=false2021.03 | 75.6 | 94.1 | 96.8 | — | — | — | — | — | — | — | — | — | — | |
| UNITERPT Dataset=CC3M, VG, COCO, SBU2022.12 | 75.6 | 94.1 | 96.8 | — | — | — | — | — | — | — | — | — | — | |
| UNITERPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 75.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UNITER2023.01 | 75.56 | 94.08 | 96.76 | 266.4 | — | — | — | — | — | — | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 75.4 | 92.5 | 95.9 | — | — | — | — | — | — | — | — | — | — | |
| SCANBackbone=ViT-224 + BERT, Patch size=14×142026.03 | 75.3 | 93.1 | — | — | — | — | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-L/14, Zero-shot=true2026.05 | 75.24 | 93.42 | 96.78 | — | — | — | — | — | — | — | — | — | — | |
| FreqAdapterFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 75.14 | 92.76 | 96.34 | — | — | — | — | — | — | — | — | — | — | |
| DreamLIPZero-shot=true2025.09 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MobileCLIP2024.06 | 74.9 | 92.9 | 96.3 | — | 82.6 | — | — | — | — | — | — | — | — | |
| ImageBindModel Size=Huge2024.06 | 74.9 | 93 | 96.1 | — | 82.7 | — | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-L/14, Zero-shot=true, Fine-tuned with=DOCCI2025.03 | 74.76 | 92.66 | — | — | — | — | — | — | — | — | 98.44 | 99.32 | — | |
| SigLIPZero-shot=true2025.09 | 74.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true2022.04 | 74.5 | 92.9 | — | — | — | — | — | — | — | — | — | — | — | |
| DivE (Ours)Backbone=ResNeXt-101 + BERT, CA=false2022.11 | 74.3 | 94 | 96.7 | — | — | — | — | — | — | — | — | — | — | |
| VSE∞Backbone=ResNeXt-101 + BERT, CA=false2022.11 | 74.2 | 93.7 | 96.8 | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-ARLEvaluation Protocol=zero-shot2024.09 | 74.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-ARLTraining Samples=20M2024.09 | 74.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GENIUSRe-ranking=true2025.03 | 74.1 | 92 | 94.8 | — | — | — | — | — | — | — | — | — | — | |
| GENIUS with Re-rankingTraining Data=Flickr30K, Zero-shot=false, Re-ranking=true2025.03 | 74.1 | 92 | 94.8 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-AdapterFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 74 | 92.84 | 96 | — | — | — | — | — | — | — | — | — | — | |
| GOALBackbone=ViT-L/14, Zero-shot=true, Fine-tuned with=DCI2025.03 | 73.76 | 91.92 | — | — | — | — | — | — | — | — | 98.22 | 99.2 | — | |
| ONE-PEACETraining Status=Pretrained2024.06 | 73.4 | 91.5 | 95.4 | — | 81.2 | — | — | — | — | — | — | — | — | |
| FFF-ViT-B/32Pre-train dataset=Open70M, Zero-shot=true2024.05 | 72.9 | 92.4 | 95.7 | — | — | — | — | — | — | — | — | — | — | |
| CLIP-L/14 + RECOBackbone=CLIP-L/14, # par.=435M, Zero-shot=true2023.06 | 72.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Align# par.=247M, Zero-shot=true2023.06 | 72.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMAFoundation Model=CLIP-B/16, Zero-shot=true2026.05 | 72.58 | 92.28 | 95.84 | — | — | — | — | — | — | — | — | — | — | |
| UNITERBackbone=R1012021.04 | 72.5 | 92.4 | 96.1 | — | — | — | — | — | — | — | — | — | — | |
| SOHOBackbone=R1012021.04 | 72.5 | 92.7 | 96.1 | — | — | — | — | — | — | — | — | — | — | |
| SOHO2026.03 | 72.5 | 92.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Fast and Slow (K=100)Object Detector=false, Input Size=384, Pre-training Data=COCO+CC, Zero-shot Setting=false2021.03 | 72.1 | 91.5 | 95.2 | — | — | — | — | — | — | — | — | — | — | |
| 3TModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 72.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FFF-ViT-B/32Pre-train dataset=Open30M, Zero-shot=true2024.05 | 72 | 91.4 | 94.9 | — | — | — | — | — | — | — | — | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Mode=Zero-shot2024.07 | 72 | 90.8 | 95 | — | — | — | — | — | — | — | — | — | — | |
| PyramidCLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true2022.04 | 71.6 | 91.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Unicoder-VLObject Detector=true, Input Size=Full, Pre-training Data=CC+SBU, Zero-shot Setting=false2021.03 | 71.5 | 90.9 | 94.9 | — | — | — | — | — | — | — | — | — | — | |
| Unicoder-VL2021.04 | 71.5 | 90.9 | 94.9 | — | — | — | — | — | — | — | — | — | — | |
| CLIPFoundation Model=CLIP-L/14-336, Zero-shot=true2026.05 | 71.4 | 91.64 | 95.46 | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-MpNetEvaluation Protocol=zero-shot2024.09 | 71.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-MpNetBackbone=MpNet, Training Samples=20M2024.09 | 71.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NegCLIPZero-shot=true2025.09 | 70.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Long-CLIPBackbone=ViT-B/16, Zero-shot=true2025.03 | 70.8 | 90.68 | — | — | — | — | — | — | — | — | 97.74 | 98.88 | — | |
| LongCLIPZero-shot=true2025.09 | 70.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 70.5 | 90.9 | — | — | — | — | — | — | — | — | — | — | — | |
| LAION-CLIPBackbone=ViT-L, Evaluation Protocol=zero-shot2024.09 | 70.2 | — | — | — | — | — | — | — | — | — | — | — | — |