Text Retrieval on Flickr30k Zero-shot (test)
94.1Recall@1ALBEF
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ALBEFPre-training volume=14M, Pre-training scale=>10M, Evaluation protocol=Zero-shot2021.11 | 94.1 | 99.5 | 99.7 | — | |
| METERBackbone=CLIP-ViT-BASE, Pre-training scale=<10M, Evaluation protocol=Zero-shot2021.11 | 90.9 | 98.3 | 99.5 | — | |
| ALBEFPre-training volume=4M, Pre-training scale=<10M, Evaluation protocol=Zero-shot2021.11 | 90.5 | 98.8 | 99.7 | — | |
| BLIP-2Backbone=ViT-L/16, Training data=400M, Storage per image=2,359 kB, Joint encoding parameters=167M, Inference time (ms)=98.642025.10 | 88.6 | — | — | — | |
| Local (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=442kB, Joint encoding parameters=33M, Inference time (ms)=4.142025.10 | 87.8 | — | — | — | |
| Compressed-128 (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=98kB, Joint encoding parameters=33M, Inference time (ms)=2.042025.10 | 87.1 | — | — | — | |
| Compressed-64 (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=49kB, Joint encoding parameters=33M, Inference time (ms)=1.912025.10 | 86.9 | — | — | — | |
| BLIPBackbone=ViT-L/16, Training data=129M, Storage per image=2,359 kB, Joint encoding parameters=139M, Inference time (ms)=101.612025.10 | 86.7 | — | — | — | |
| METERBackbone=Swin-BASE, Pre-training scale=<10M, Evaluation protocol=Zero-shot2021.11 | 85.3 | 97.7 | 99.2 | — | |
| BLIPBackbone=ViT-B/16, Training data=12M, Storage per image=1,769 kB, Joint encoding parameters=139M, Inference time (ms)=83.272025.10 | 84.9 | — | — | — | |
| Local (EDJE)Backbone=ViT-B/16, Training data=12M, Storage per image=442kB, Joint encoding parameters=33M, Inference time (ms)=2.862025.10 | 84.3 | — | — | — | |
| UNITER_LARGEPre-training scale=<10M, Evaluation protocol=Zero-shot2021.11 | 83.6 | 95.7 | 97.7 | — | |
| ALBEFBackbone=ViT-B/16, Training data=12M, Storage per image=1,769 kB, Joint encoding parameters=147M, Inference time (ms)=45.922025.10 | 82.8 | — | — | — | |
| 3SHNetVisual Representation=Region+Grid, Ensemble=false2024.04 | 74.9 | 93.2 | 96.2 | 490 | |
| ViLTPre-training scale=<10M, Evaluation protocol=Zero-shot2021.11 | 73.2 | 93.6 | 96.5 | — | |
| 3SHNetVisual Representation=Region, Ensemble=true2024.04 | 72.7 | 90.9 | 94.2 | 478.6 | |
| 3SHNetVisual Representation=Grid, Ensemble=false2024.04 | 70.9 | 91.6 | 94.9 | 477.8 | |
| ESAEnsemble=false2024.04 | 69.5 | 89.1 | 93.8 | 467.4 | |
| VSE∞Ensemble=false, Source=derived from released pre-trained model2024.04 | 68 | 89.2 | 93.7 | 462.8 | |
| DIMEEnsemble=true, Source=derived from released pre-trained model2024.04 | 67.4 | 90.1 | 94.5 | 471.4 | |
| SGRAFEnsemble=true2024.04 | 65.7 | 87.2 | 93.4 | 450.2 | |
| DIMEEnsemble=false, Source=derived from released pre-trained model2024.04 | 63.5 | 86.9 | 93.1 | 453.2 | |
| CLIP-PGSEvaluation Protocol=Zero-shot, Lower limit masking rate=0.32025.03 | 59.9 | 83.5 | 90.8 | — | |
| CLIPEvaluation Protocol=Zero-shot2025.03 | 58.5 | 83.8 | 89.1 | — | |
| CLIP-PGSEvaluation Protocol=Zero-shot, Lower limit masking rate=0.52025.03 | 57.7 | 82.7 | 90.4 | — | |
| CVSEEnsemble=false2024.04 | 56.4 | 83 | 89 | 414.1 | |
| E-CLIPEvaluation Protocol=Zero-shot2025.03 | 55.8 | 84.2 | 89.6 | — | |
| A-CLIPEvaluation Protocol=Zero-shot2025.03 | 55.3 | 81.4 | 87.6 | — | |
| FLIPEvaluation Protocol=Zero-shot2025.03 | 53.8 | 80.8 | 88.5 | — | |
| SGREnsemble=false2024.04 | 51.4 | 79.2 | 87.2 | 404.6 |