Text-to-Image Retrieval on MS-COCO (test)
2,208R@1K-LITE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| K-LITETraining Data Dataset=GCC-15M + ImageNet-21K, Training Data # Samples=15M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 2,208 | 45.41 | — | |
| UniCLTraining Data Dataset=GCC-15M + ImageNet-21K, Training Data # Samples=15M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 2,141 | 44.69 | — | |
| K-LITETraining Data Dataset=YFCC-14M + ImageNet-21K, Training Data # Samples=14M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 1,438 | 33.77 | — | |
| UniCLTraining Data Dataset=YFCC-14M + ImageNet-21K, Training Data # Samples=14M (half), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 1,333 | 32.14 | — | |
| K-LITETraining Data Dataset=ImageNet-21K, Training Data # Samples=13M (full), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 191 | 6.48 | — | |
| UniCLTraining Data Dataset=ImageNet-21K, Training Data # Samples=13M (full), Evaluation Protocol=zero-shot, Pre-training Epochs=322022.04 | 98 | 3.54 | — | |
| BLIP-2Throughput=1.68/s2024.06 | 66.3 | 86.5 | 91.8 | |
| BLIPPre-train # Images=14M2023.12 | 63.1 | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Clean2023.12 | 61.6 | — | — | |
| MAFAPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 61.2 | — | — | |
| ALBEFPre-train # Images=14M2023.12 | 60.7 | — | — | |
| GRIT-VLPPre-train # Images=4M, Pre-train Dataset=4M-Noisy, Reproduced=true2023.12 | 59.6 | — | — | |
| TCLPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 59 | — | — | |
| BLIPPre-train # Images=4M, Pre-train Dataset=4M-Clean, Reproduced=true2023.12 | 58.9 | — | — | |
| InternVL-GThroughput=2.03/s2024.06 | 58.6 | 81.3 | 88 | |
| ALBEFPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 56.8 | — | — | |
| OSCARPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 54 | — | — | |
| UNITERPre-train # Images=4M, Pre-train Dataset=4M-Noisy2023.12 | 52.9 | — | — | |
| CARTThroughput=105.8/s2024.06 | 52.4 | 77.5 | 86.1 | |
| 3TModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 48.5 | — | — | |
| FIMA-Q w/ RegCacheBackbone=SigLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 47.64 | — | — | |
| Naïve w/ RegCacheBackbone=SigLIP-B/16, Quantization (W/A bits)=W8A8, Zero-shot=true2025.10 | 46.3 | — | — | |
| GENIUSRe-ranking=true2025.03 | 46.1 | 74 | 82.7 | |
| Baseline (CLIP-style)Model scale=g scale, Training dataset=Text-Filtered WebLI2023.05 | 44.7 | — | — | |
| LiTModel scale=g scale, Training dataset=Text-Filtered WebLI, Pre-training dataset=JFT2023.05 | 43.6 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true2022.04 | 42.6 | 68.6 | — | |
| NaïveBackbone=SigLIP-B/16, Quantization (W/A bits)=W8A8, Zero-shot=true2025.10 | 41.8 | — | — | |
| RCSR-pMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 40.38 | 68.35 | — | |
| Align# par.=247M, Zero-shot=true2023.06 | 40.2 | — | — | |
| GENIUSRe-ranking=false2025.03 | 40.1 | 66.2 | 75.8 | |
| PyramidCLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true2022.04 | 39.6 | 66.2 | — | |
| RCSRMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 39.46 | 66.92 | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset=143M, Zero-shot=true2022.04 | 38.8 | 64.9 | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 38.8 | 65 | — | |
| CLIP-L/14 + RECOBackbone=CLIP-L/14, # par.=435M, Zero-shot=true2023.06 | 38.7 | — | — | |
| CreamFLMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 38.58 | 66.18 | — | |
| Flava# par.=172M, Zero-shot=true2023.06 | 38.4 | — | — | |
| pFedMMAMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 38.17 | 66.25 | — | |
| FedProxMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 37.22 | 64.96 | — | |
| MFCPLMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 36.68 | 64.48 | — | |
| FedAvgMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=false2026.04 | 36.14 | 63.84 | — | |
| FedMEKTMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 35.92 | 63.58 | — | |
| MOONMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 35.24 | 62.94 | — | |
| CLIP-L/14Backbone=CLIP-L/14, # par.=428M, Zero-shot=true2023.06 | 35.2 | — | — | |
| CLIPImage Encoder=ViT-B/32, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 34.4 | 60.6 | — | |
| CLIPImage Encoder=ResNet50, Pretrain Dataset=143M, Zero-shot=true, Variant=Our implementation2022.04 | 34 | 60 | — | |
| FedPerMissing Modality Rate=50%, Client Participation Rate=50%, Number of Clients=30, Communication Rounds=200, alpha (Dirichlet coefficient)=0.1, Personalization=true2026.04 | 33.68 | 61.82 | — | |
| CLIP-B/32 + RECOBackbone=CLIP-B/32, # par.=154M, Zero-shot=true2023.06 | 33.6 | — | — | |
| FIMA-Q w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W4A4, Zero-shot=true2025.10 | 33.43 | — | — | |
| FP32Backbone=CLIP-B/16, Quantization (W/A bits)=FP32, Zero-shot=true2025.10 | 32.73 | — | — | |
| CLIPImage Encoder=ViT-B/16, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 32.5 | 57.7 | — | |
| CLIP-B/32Backbone=CLIP-B/32, # par.=151M, Zero-shot=true2023.06 | 30.2 | — | — | |
| CLIPImage Encoder=ViT-B/32, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 29.8 | 54.4 | — | |
| IRGenRe-ranking=false2025.03 | 29.6 | 50.7 | 56.3 | |
| Naïve w/ RegCacheBackbone=CLIP-B/16, Quantization (W/A bits)=W8A8, Zero-shot=true2025.10 | 28.02 | — | — | |
| CLIPImage Encoder=ResNet50, Pretrain Dataset=400M, Zero-shot=true, Variant=Released model2022.04 | 27.4 | 51.8 | — | |
| DeCLIPImage Encoder=ViT-B/32, Pretrain Dataset=88M, Zero-shot=true, Variant=Pretrained model w text encoder2022.04 | 22.1 | 45.8 | — | |
| DeCLIPImage Encoder=ResNet50, Pretrain Dataset=88M, Zero-shot=true, Variant=Pretrained model w text encoder2022.04 | 21.7 | 44.6 | — | |
| Ours2025.12 | 16.79 | 33.6 | — | |
| GRACERe-ranking=false2025.03 | 16.7 | 39.2 | 50.3 | |
| ECLIP2025.12 | 15.76 | 32.03 | — | |
| B-cosified2025.12 | 15.58 | 30.89 | — | |
| GAME2025.12 | 15.37 | 30.83 | — | |
| M2IB2025.12 | 14.69 | 30.04 | — | |
| MaskCLIP2025.12 | 14.23 | 29.53 | — | |
| CLIPSurgery2025.12 | 14.19 | 29.41 | — | |
| Rollout2025.12 | 12.94 | 29.32 | — | |
| Full DatasetEvaluation Model=ResNet, Pairs=Full Dataset2026.02 | 12.6 | 33.4 | 46.5 | |
| Full DatasetEvaluation Model=ViT, Pairs=Full Dataset2026.02 | 11.5 | 29.8 | 41.7 | |
| Grad-CAM2025.12 | 10.27 | 22.16 | — | |
| PDSEvaluation Model=ResNet, Pairs=3002026.02 | 5.3 | 17.2 | 27.2 | |
| PDSEvaluation Model=ViT, Pairs=3002026.02 | 4.1 | 13.4 | 21.2 | |
| TESLA-VLEvaluation Model=ResNet, Pairs=3002026.02 | 3 | 10.7 | 17.6 | |
| PDSEvaluation Model=ResNet, Pairs=1002026.02 | 2.8 | 10 | 17.3 | |
| LoRSEvaluation Model=ResNet, Pairs=3002026.02 | 2.5 | 8.5 | 13.8 | |
| PDSEvaluation Model=ViT, Pairs=1002026.02 | 2.3 | 8.6 | 14.5 | |
| LoRSEvaluation Model=ResNet, Pairs=1002026.02 | 1.8 | 6.8 | 11.4 | |
| TESLA-VLEvaluation Model=ViT, Pairs=3002026.02 | 1.5 | 5.9 | 10.1 | |
| TESLA-VLEvaluation Model=ResNet, Pairs=1002026.02 | 1.4 | 5.8 | 10.2 | |
| LoRSEvaluation Model=ViT, Pairs=3002026.02 | 1 | 3.7 | 6.3 | |
| LoRSEvaluation Model=ViT, Pairs=1002026.02 | 0.8 | 3 | 5.4 | |
| TESLA-VLEvaluation Model=ViT, Pairs=1002026.02 | 0.5 | 2.1 | 3.8 | |
| CLIPPretrain=CC3M, Evaluation Protocol=Zero-shot2024.04 | — | 23 | — | |
| CLIPPretrain=CC12M, Evaluation Protocol=Zero-shot2024.04 | — | 39 | — | |
| CLIPPretrain=DataComp, Evaluation Protocol=Zero-shot2024.04 | — | 16 | — | |
| Codebook-CLIPPretrain=CC3M, Evaluation Protocol=Zero-shot2024.04 | — | 28 | — | |
| Codebook-CLIPPretrain=CC12M, Evaluation Protocol=Zero-shot2024.04 | — | 45 | — | |
| Codebook-CLIPPretrain=DataComp, Evaluation Protocol=Zero-shot2024.04 | — | 20 | — | |
| IL-CLIPPretrain=CC3M, Evaluation Protocol=Zero-shot2024.04 | — | 28 | — | |
| IL-CLIPPretrain=CC12M, Evaluation Protocol=Zero-shot2024.04 | — | 44 | — | |
| IL-CLIPPretrain=DataComp, Evaluation Protocol=Zero-shot2024.04 | — | 18 | — | |
| NegCLIPPretrain=CC3M, Evaluation Protocol=Zero-shot2024.04 | — | 19 | — | |
| NegCLIPPretrain=CC12M, Evaluation Protocol=Zero-shot2024.04 | — | 36 | — | |
| NegCLIPPretrain=DataComp, Evaluation Protocol=Zero-shot2024.04 | — | 13 | — |