Image-to-text Retrieval on MS-COCO (test)
80.6R@1BLIP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| BLIPPre-train # Images=14M2023.12 | 80.6 | — | — | |
| L2RM-SGRAFMRate=0.22024.03 | 80.2 | 96.3 | 98.5 | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 79.4 | — | — | |
| RCL-SGRAFMRate=0.22024.03 | 78.9 | 96 | 98.4 | |
| L2RM-SGRMRate=0.22024.03 | 78.4 | 95.7 | 98.3 | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 78 | — | — | |
| L2RM-SAFMRate=0.22024.03 | 77.9 | 96 | 98.3 | |
| ALBEFPre-train # Images=14M2023.12 | 77.6 | — | — | |
| DECL-SGRAFMRate=0.22024.03 | 77.5 | 95.9 | 98.4 | |
| L2RM-SGRAFMRate=0.42024.03 | 77.5 | 95.8 | 98.4 | |
| RCL-SGRMRate=0.22024.03 | 77 | 95.5 | 98.1 | |
| RCL-SGRAFMRate=0.42024.03 | 77 | 95.5 | 98.3 | |
| NCRMRate=0.22024.03 | 76.6 | 95.6 | 98.2 | |
| BiCroMRate=0.22024.03 | 76.6 | 95.4 | 98.2 | |
| GRIT-VLPPre-train # Images=4M, Pre-train Dataset=4M-Noisy, Reproduced=true2023.12 | 76.6 | — | — | |
| DECL-SGRMRate=0.22024.03 | 75.6 | 95.1 | 98.3 | |
| DECL-SGRAFMRate=0.42024.03 | 75.6 | 95.5 | 98.3 | |
| TCLPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 75.6 | — | — | |
| BLIPPre-train # Images=4M, Pre-train Dataset=4M-Clean, Reproduced=true2023.12 | 75.5 | — | — | |
| L2RM-SGRAFMRate=0.62024.03 | 75.4 | 94.7 | 97.9 | |
| BiCroMRate=0.42024.03 | 75.2 | 95.3 | 98.1 | |
| L2RM-SGRMRate=0.42024.03 | 75.2 | 94.8 | 98.1 | |
| NCRMRate=0.42024.03 | 74.7 | 94.6 | 98 | |
| L2RM-SAFMRate=0.42024.03 | 74.4 | 94.7 | 98.3 | |
| RCL-SGRAFMRate=0.62024.03 | 74 | 94.3 | 97.5 | |
| RCL-SGRMRate=0.42024.03 | 73.9 | 94.9 | 97.9 | |
| DECL-SGRMRate=0.42024.03 | 73.6 | 94.6 | 97.9 | |
| BiCroMRate=0.62024.03 | 73.2 | 93.9 | 97.6 | |
| ALBEFPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 73.1 | — | — | |
| DECL-SGRAFMRate=0.62024.03 | 73 | 94.2 | 97.9 | |
| L2RM-SGRMRate=0.62024.03 | 72.7 | 93.9 | 97.5 | |
| RCL-SGRMRate=0.62024.03 | 71.4 | 93.2 | 97.1 | |
| L2RM-SAFMRate=0.62024.03 | 71.2 | 93.4 | 97.5 | |
| OSCARPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 70 | — | — | |
| IMRAMMRate=0.22024.03 | 69.9 | 93.6 | 97.4 | |
| DECL-SGRMRate=0.62024.03 | 69.7 | 93.4 | 97.5 | |
| L2RM-SGRAFMRate=0.82024.03 | 69 | 91.9 | 96.4 | |
| FIMA-Q w/ RegCacheBackbone=SigLIP-B/16, Quantization (W/A bits)=W8A8, Zero-shot=true2025.10 | 68.18 | — | — | |
| FIMA-Q w/ RegCacheBackbone=SigLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 68.14 | — | — | |
| FIMA-Q w/ RegCacheBackbone=SigLIP-B/16, Quantization (W/A bits)=W6A6, Zero-shot=true2025.10 | 68.14 | — | — | |
| RCL-SGRAFMRate=0.82024.03 | 67.4 | 90.8 | 96 | |
| UNITERPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 65.7 | — | — | |
| L2RM-SGRMRate=0.82024.03 | 65.2 | 90.3 | 96.1 | |
| DECL-SGRAFMRate=0.82024.03 | 64.8 | 90.5 | 96 | |
| L2RM-SAFMRate=0.82024.03 | 64.7 | 90.8 | 95.8 | |
| 3TModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 64.1 | — | — | |
| RCL-SGRMRate=0.82024.03 | 63.2 | 89.3 | 95.2 | |
| BiCroMRate=0.82024.03 | 62.2 | 88.6 | 94.6 | |
| Baseline (CLIP-style)Model scale=g scale, Training dataset=Text-Filtered WebLI2023.05 | 60 | — | — | |
| DECL-SGRMRate=0.82024.03 | 60 | 88.7 | 94.5 | |
| LiTModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 59.5 | — | — | |
| CLIP-L/14 + RECOBackbone=CLIP-L/14, # par.=435M, Zero-shot=true2023.06 | 58 | — | — | |
| CLIP-L/14Backbone=CLIP-L/14, # par.=428M, Zero-shot=true2023.06 | 57.2 | — | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 56.9 | 79.6 | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true2022.04 | 55.7 | 80.8 | — | |
| Align# par.=247M, Zero-shot=true2023.06 | 55.1 | — | — | |
| PyramidCLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true2022.04 | 55 | 79.8 | — | |
| FIMA-Q w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W8A8, Zero-shot=true2025.10 | 53.66 | — | — | |
| FIMA-Q w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W6A6, Zero-shot=true2025.10 | 53.22 | — | — | |
| FP32Backbone=CLIP-B/16, Quantization (W/A bits)=FP32, Zero-shot=true2025.10 | 52.94 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset=143M, Zero-shot=true2022.04 | 52.8 | 78.1 | — | |
| CLIP-B/32 + RECOBackbone=CLIP-B/32, # par.=154M, Zero-shot=true2023.06 | 52.2 | — | — | |
| FIMA-Q w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 51.82 | — | — | |
| CLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 51.8 | 76.4 | — | |
| IMRAMMRate=0.42024.03 | 51.8 | 82.4 | 90.9 | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 51.7 | 76.1 | — | |
| CLIP-B/32Backbone=CLIP-B/32, # par.=151M, Zero-shot=true2023.06 | 51.2 | — | — | |
| CLIPImage Encoder=ViT-B/32, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 51.1 | 76.4 | — | |
| CLIPImage Encoder=ViT-B/32, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 49.2 | 74.1 | — | |
| CLIPImage Encoder=ResNet50, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 47.6 | 73.1 | — | |
| Flava# par.=172M, Zero-shot=true2023.06 | 42.7 | — | — | |
| FIMA-QBackbone=CLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 40.92 | — | — | |
| RCSR-pMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 39.12 | 67.48 | — | |
| RCSRMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 38.26 | 66.24 | — | |
| CreamFLMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 37.42 | 65.86 | — | |
| pFedMMAMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 36.95 | 64.85 | — | |
| FedProxMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 36.18 | 64.52 | — | |
| MFCPLMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 35.84 | 64.12 | — | |
| FedAvgMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 35.26 | 63.68 | — | |
| FedMEKTMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 34.86 | 63.12 | — | |
| MOONMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 34.12 | 62.46 | — | |
| FedPerMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 32.84 | 61.24 | — | |
| K-LITETraining Data Dataset=GCC-15M + ImageNet-21K, Training Data # Samples=15M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 32.68 | 58.88 | — | |
| DeCLIPImage Encoder=ViT-B/32, Pretrain Dataset=88M, Zero-shot=true, Variant=Pretrained model w text encoder2022.04 | 32.6 | 59.1 | — | |
| DeCLIPImage Encoder=ResNet50, Pretrain Dataset=88M, Zero-shot=true, Variant=Pretrained model w text encoder2022.04 | 32 | 57.8 | — | |
| UniCLTraining Data Dataset=GCC-15M + ImageNet-21K, Training Data # Samples=15M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 31.88 | 57.76 | — | |
| Full DatasetEvaluation Model=ResNet, Pairs=Full Dataset2026.02 | 26.6 | 53.1 | 66.1 | |
| K-LITETraining Data Dataset=YFCC-14M + ImageNet-21K, Training Data # Samples=14M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 22.44 | 47.28 | — | |
| Ours2025.12 | 21.83 | 39.02 | — | |
| UniCLTraining Data Dataset=YFCC-14M + ImageNet-21K, Training Data # Samples=14M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 21.8 | 45.38 | — | |
| GAME2025.12 | 20.97 | 37.35 | — | |
| M2IB2025.12 | 20.58 | 36.91 | — | |
| ECLIP2025.12 | 20.56 | 37.61 | — | |
| B-cosified2025.12 | 20.16 | 36.22 | — | |
| Full DatasetEvaluation Model=ViT, Pairs=Full Dataset2026.02 | 19.5 | 42.4 | 55.3 | |
| MaskCLIP2025.12 | 18.91 | 35.14 | — | |
| IMRAMMRate=0.62024.03 | 18.2 | 51.6 | 68 | |
| CLIPSurgery2025.12 | 17.71 | 33.84 | — | |
| Rollout2025.12 | 17.53 | 35.03 | — | |
| ERQ w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 13.78 | — | — |