Image-text retrieval on Flickr30K (test)
94.1R@1 (Img->Txt)Uni-Perceiver-L + Conditional MoEs
Evaluation Results
| Method | Links | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Uni-Perceiver-L + Conditional MoEs#param=505M, #data=44.1M, tuning=fine-tuning (100%)2022.06 | 94.1 | — | — | — | — | — | 83.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, tuning=fine-tuning (100%)2022.06 | 93.6 | — | — | — | — | — | 79.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FIBER-Rerank-50strategy=re-ranking, top-k=502022.06 | 91.08 | — | — | — | — | — | 96.1 | 98.5 | 99.37 | 99.7 | 100 | — | — | — | — | — | — | — | — | |
| FIBER-Rerank-100strategy=re-ranking, top-k=1002022.06 | 91.02 | — | — | — | — | — | 96 | 98.54 | 99.34 | 99.7 | 100 | — | — | — | — | — | — | — | — | |
| FIBER-ITC+ITM Ensemblestrategy=ensemble2022.06 | 90.96 | — | — | — | — | — | 96 | 98.44 | 99.14 | 99.7 | 100 | — | — | — | — | — | — | — | — | |
| FIBER-Rerank-10strategy=re-ranking, top-k=102022.06 | 90.94 | — | — | — | — | — | 95.8 | 98.16 | 98.48 | 99.6 | 99.9 | — | — | — | — | — | — | — | — | |
| FIBER-Rerank-20strategy=re-ranking, top-k=202022.06 | 90.1 | — | — | — | — | — | 95.9 | 98.38 | 99.14 | 99.8 | 100 | — | — | — | — | — | — | — | — | |
| BLIP_CapFilt-LPretrain Images=129M, Retrieval Strategy=re-ranking2022.06 | 87.5 | — | — | — | — | — | 97.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| X-VLMPretrain Images=4M, Retrieval Strategy=re-ranking2022.06 | 86.9 | — | — | — | — | — | 97 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CDDSBackbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 86.8 | — | — | — | — | — | 76.3 | 98.3 | 99.6 | 94.3 | 97.2 | 552.5 | — | — | — | — | — | — | — | |
| X-VLM2022.06 | 86.1 | — | — | — | — | — | 96.8 | 97.4 | 98.7 | 99.8 | 100 | — | — | — | — | — | — | — | — | |
| LAPSBackbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 85.1 | — | — | — | — | — | 74 | 97.7 | 99.2 | 93 | 96.3 | 545.3 | — | — | — | — | — | — | — | |
| VLMo-LPretrain Images=4M, Retrieval Strategy=dual encoder2022.06 | 84.5 | — | — | — | — | — | 95.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FIBER-ITMstrategy=fusion-encoder2022.06 | 84.1 | — | — | — | — | — | 95.1 | 97.54 | 98.88 | 99.6 | 99.9 | — | — | — | — | — | — | — | — | |
| VSE++Backbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 83.3 | — | — | — | — | — | 71.1 | 97.5 | 99.2 | 93.2 | 96.2 | 540.6 | — | — | — | — | — | — | — | |
| CDDSBackbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 83 | — | — | — | — | — | 70.3 | 98.1 | 99.7 | 92.1 | 95.8 | 539 | — | — | — | — | — | — | — | |
| ALBEF2022.06 | 82.8 | — | — | — | — | — | 94.3 | 96.7 | 98.4 | 99.4 | 99.8 | — | — | — | — | — | — | — | — | |
| ALBEF-BPretrain Images=4M, Retrieval Strategy=re-ranking2022.06 | 82.8 | — | — | — | — | — | 94.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VSE++Backbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 82.5 | — | — | — | — | — | 70 | 96.5 | 98.9 | 91.4 | 95.1 | 534.4 | — | — | — | — | — | — | — | |
| LAPSBackbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 82.4 | — | — | — | — | — | 70 | 97.4 | 99.5 | 91.7 | 95.4 | 536.3 | — | — | — | — | — | — | — | |
| SCANBackbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 81.9 | — | — | — | — | — | 70 | 96.9 | 98.9 | 92.7 | 95.8 | 536.1 | — | — | — | — | — | — | — | |
| FIBER-ITCstrategy=dual-encoder2022.06 | 81.44 | — | — | — | — | — | 92.9 | 96.72 | 98.48 | 99.5 | 99.9 | — | — | — | — | — | — | — | — | |
| FIBER-BPretrain Images=4M, Retrieval Strategy=dual encoder2022.06 | 81.44 | — | — | — | — | — | 92.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CHANBackbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 81.4 | — | — | — | — | — | 68.5 | 97 | 98.6 | 90.6 | 94.5 | 530.6 | — | — | — | — | — | — | — | |
| CHANBackbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 81.2 | — | — | — | — | — | 70.3 | 96.7 | 98.8 | 92.2 | 95.9 | 535 | — | — | — | — | — | — | — | |
| SGRBackbone=Swin, Resolution=384x384, Number of patches=12x122026.03 | 80.7 | — | — | — | — | — | 69.9 | 96.8 | 99 | 91.7 | 95.3 | 533.4 | — | — | — | — | — | — | — | |
| SGRBackbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 80.4 | — | — | — | — | — | 66.9 | 97 | 98.7 | 90.2 | 94.5 | 527.6 | — | — | — | — | — | — | — | |
| CDDSBackbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 79.5 | — | — | — | — | — | 67.5 | 96.3 | 98.1 | 90.8 | 94.6 | 526.8 | — | — | — | — | — | — | — | |
| VLMo-BPretrain Images=4M, Retrieval Strategy=dual encoder2022.06 | 79.3 | — | — | — | — | — | 92.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| METER-Swin-BPretrain Images=4M, Retrieval Strategy=fusion encoder2022.06 | 79.02 | — | — | — | — | — | 92.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 79 | — | — | — | — | — | 67.3 | 96 | 98.1 | 90.5 | 94.5 | 525.4 | — | — | — | — | — | — | — | |
| SCANBackbone=Swin, Resolution=224x224, Number of patches=7x72026.03 | 79 | — | — | — | — | — | 67.7 | 95.9 | 98.2 | 90.6 | 94.9 | 526.3 | — | — | — | — | — | — | — | |
| VSE++Backbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 77.1 | — | — | — | — | — | 65.8 | 95.7 | 97.5 | 90.2 | 94.3 | 520.5 | — | — | — | — | — | — | — | |
| SGRBackbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 76.9 | — | — | — | — | — | 64.2 | 94.9 | 98.1 | 88.4 | 93.3 | 515.8 | — | — | — | — | — | — | — | |
| SCANBackbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 75.4 | — | — | — | — | — | 63.6 | 94.4 | 96.9 | 88.6 | 93.5 | 512.5 | — | — | — | — | — | — | — | |
| CHANBackbone=ViT, Resolution=384x384, Number of patches=24x242026.03 | 75.4 | — | — | — | — | — | 63.2 | 94.5 | 97.6 | 88.6 | 93.1 | 512.4 | — | — | — | — | — | — | — | |
| CDDSBackbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 74.8 | — | — | — | — | — | 63.1 | 93.6 | 97.8 | 88.2 | 93.1 | 510.6 | — | — | — | — | — | — | — | |
| VILLA-BPretrain Images=4M, Retrieval Strategy=fusion encoder2022.06 | 74.7 | — | — | — | — | — | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LAPSBackbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 74 | — | — | — | — | — | 62.5 | 93.4 | 97.4 | 87.3 | 92.7 | 507.3 | — | — | — | — | — | — | — | |
| UNITER-BPretrain Images=4M, Retrieval Strategy=fusion encoder2022.06 | 72.5 | — | — | — | — | — | 85.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VSE++Backbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 71.8 | — | — | — | — | — | 59.4 | 92.8 | 96.5 | 84.7 | 90.9 | 496.1 | — | — | — | — | — | — | — | |
| SGRBackbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 69.7 | — | — | — | — | — | 59.1 | 90.8 | 95.2 | 84.1 | 89.9 | 488.7 | — | — | — | — | — | — | — | |
| SCANBackbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 69.5 | — | — | — | — | — | 56.4 | 90.9 | 95.6 | 83.1 | 90 | 485.6 | — | — | — | — | — | — | — | |
| CHANBackbone=ViT, Resolution=224x224, Number of patches=14x142026.03 | 69.2 | — | — | — | — | — | 58.4 | 91.8 | 95 | 84.9 | 90.6 | 489.9 | — | — | — | — | — | — | — | |
| ViLT-BPretrain Images=4M, Retrieval Strategy=fusion encoder2022.06 | 64.4 | — | — | — | — | — | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ALBEF2023.01 | — | — | — | — | — | — | — | — | — | — | — | 564.58 | 0 | — | — | — | — | — | — | |
| B12023.01 | — | — | — | — | — | — | — | — | — | — | — | 559.82 | -4.76 | — | — | — | — | — | — | |
| B22023.01 | — | — | — | — | — | — | — | — | — | — | — | 563.04 | -1.54 | — | — | — | — | — | — | |
| BadCLIPtraining_data=3%2023.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | 85.2 | 98.3 | — | — | — | — | |
| C2LIPMethod Category=Models fine-tuned by us2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 97 | — | — | — | — | — | |
| CCLMMultilingual Multimodal Pretraining=true2022.11 | — | 96 | 93.3 | 93.7 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CE-CLIPMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 87.4 | — | — | — | — | — | |
| CLICMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.1 | — | — | — | — | — | |
| CLIP2023.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | 83 | — | — | — | — | — | |
| CLIPEvaluation Protocol=zero-shot, Venue=ICML’21, Data=CLIP 400M2024.06 | — | — | — | — | — | — | — | — | — | — | — | 540.6 | — | — | — | — | — | — | — | |
| CLIP (OpenAI)Backbone=ViT-B/322026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 89 | — | — | — | — | — | |
| CLIP (OpenAI)Backbone=ViT-B/162026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.9 | — | — | — | — | — | |
| CoCoOptraining_data=3%2023.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | 85.9 | — | — | — | — | — | |
| Codebook-CLIPMethod Category=Codebook-based models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.6 | — | — | — | — | — | |
| CoN-CLIPMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 86 | — | — | — | — | — | |
| CoOptraining_data=3%2023.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | 79.4 | — | — | — | — | — | |
| CORA-BERTData=CC1M2024.06 | — | — | — | — | — | — | — | — | — | — | — | 530.1 | — | — | — | — | — | — | — | |
| CovMatch# Pairs=5002026.06 | — | — | — | — | — | 28.9 | — | — | — | — | — | — | — | — | — | — | — | 26.3 | 31.5 | |
| CovMatch# Pairs=10002026.06 | — | — | — | — | — | 28.4 | — | — | — | — | — | — | — | — | — | — | — | 26.5 | 30.2 | |
| DAC-LLMMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 83.7 | — | — | — | — | — | |
| DAC-SAMMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 84 | — | — | — | — | — | |
| DreamLIP-3mMethod Category=Fine-grained models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 83.1 | — | — | — | — | — | |
| EDGE# Pairs=5002026.06 | — | — | — | — | — | 25.8 | — | — | — | — | — | — | — | — | — | — | — | 19.4 | 32.1 | |
| EDGE# Pairs=10002026.06 | — | — | — | — | — | 30.5 | — | — | — | — | — | — | — | — | — | — | — | 26.2 | 34.8 | |
| EmbNFine-tuning Strategy=SOTA without pre-training2021.04 | — | 72 | 60.3 | 54.8 | 46.3 | 65.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FG-CLIPMethod Category=Fine-grained models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.8 | — | — | — | — | — | |
| FineCLIPMethod Category=Fine-grained models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 93.5 | — | — | — | — | — | |
| FLAIR-3mMethod Category=Fine-grained models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.4 | — | — | — | — | — | |
| HADA2023.01 | — | — | — | — | — | — | — | — | — | — | — | 568.22 | 3.64 | — | — | — | — | — | — | |
| IL-CLIPMethod Category=Codebook-based models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.4 | — | — | — | — | — | |
| LightningDOTImplementation=Re-implementation2023.01 | — | — | — | — | — | — | — | — | — | — | — | 532.26 | -32.32 | — | — | — | — | — | — | |
| LightningDOTImplementation=Originally reported2023.01 | — | — | — | — | — | — | — | — | — | — | — | 535.9 | -28.68 | — | — | — | — | — | — | |
| LLIPMethod Category=Fine-grained models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.2 | — | — | — | — | — | |
| M3PFine-tuning Strategy=English-only Fine-tune2021.04 | — | 87.4 | 58.5 | 46 | 36.8 | 60.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| M3PFine-tuning Strategy=Single-Language Fine-tune2021.04 | — | 87.4 | 82.1 | 67.3 | 65 | 78 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| M3PFine-tuning Strategy=All-Language Fine-tune2021.04 | — | 87.7 | 82.7 | 73.9 | 72.2 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| M3PMultilingual Multimodal Pretraining=true2022.11 | — | 87.7 | 82.7 | 73.9 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MULEFine-tuning Strategy=SOTA without pre-training2021.04 | — | 70.3 | 64.1 | 62.3 | 57.7 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MURALMultilingual Multimodal Pretraining=true, Size=base2022.11 | — | 92.2 | 88.6 | 87.6 | 84.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MURALMultilingual Multimodal Pretraining=true, Size=large2022.11 | — | 93.8 | 90.4 | 89.9 | 87.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NegCLIPMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.4 | — | — | — | — | — | |
| PAR.EmbNFine-tuning Strategy=SOTA without pre-training2021.04 | — | 69 | 62.6 | 60.6 | 54.1 | 67.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RAHA# Pairs=5002026.06 | — | — | — | — | — | 32.9 | — | — | — | — | — | — | — | — | — | — | — | 30 | 35.9 | |
| RAHA# Pairs=10002026.06 | — | — | — | — | — | 38 | — | — | — | — | — | — | — | — | — | — | — | 34.5 | 41.5 | |
| S-LIWEFine-tuning Strategy=SOTA without pre-training2021.04 | — | 76.3 | 72.1 | 63.4 | 59.4 | 70.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SigLIPBackbone=ViT-B/162026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.2 | — | — | — | — | — | |
| SigLIPBackbone=ViT-B/16, Fine-tuning Dataset=CC3M, Method Category=Models fine-tuned by us2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.6 | — | — | — | — | — | |
| SLVC-RMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.1 | — | — | — | — | — | |
| SLVC-RLMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 90 | — | — | — | — | — | |
| SMALRFine-tuning Strategy=SOTA without pre-training2021.04 | — | 74.5 | 69.8 | 65.9 | 64.8 | 73 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TripletCLIPMethod Category=Composition-aware models2026.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | 81.7 | — | — | — | — | — | |
| UC2Fine-tuning Strategy=English-only Fine-tune2021.04 | — | 87.2 | 74.9 | 74 | 67.9 | 78 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UC2Fine-tuning Strategy=Single-Language Fine-tune2021.04 | — | 87.2 | 83.8 | 77.6 | 74.2 | 83.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UC2Fine-tuning Strategy=All-Language Fine-tune2021.04 | — | 88.2 | 84.5 | 83.9 | 81.2 | 86.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UC2Multilingual Multimodal Pretraining=true2022.11 | — | 88.2 | 84.5 | 83.9 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |