Text-to-Image Retrieval on Flickr8k
57R@1Ours-Embedding
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ours-Embeddingstrategy=Embedding, mode=Zero-shot2024.05 | 57 | 82.7 | 90.1 | |
| CLIPmode=Zero-shot2024.05 | 55.7 | 81.6 | 89.9 | |
| Ours-Kmeansstrategy=Kmeans, mode=Zero-shot2024.05 | 55.1 | 83.3 | 90.9 | |
| FLIPmode=Zero-shot2024.05 | 55 | 80.9 | 88.9 | |
| Ours-RGB0.3strategy=RGB, anchor_patch_ratio=0.3, mode=Zero-shot2024.05 | 54.7 | 81.4 | 90.7 | |
| FLIPAttnmode=Zero-shot2024.05 | 53.7 | 80.09 | 87.99 | |
| Ours-RGB0.5strategy=RGB, anchor_patch_ratio=0.5, mode=Zero-shot2024.05 | 52.1 | 81.2 | 89.5 | |
| CLIP-RefineBackbone=ViT-B/32, Zero-shot=true2025.04 | 36.11 | 61.29 | 70.51 | |
| HyCDBackbone=ViT-B/32, Zero-shot=true2025.04 | 35.72 | 60.34 | 70.14 | |
| m²-mixBackbone=ViT-B/32, Zero-shot=true2025.04 | 35.04 | 59.66 | 69.66 | |
| ContrastiveBackbone=ViT-B/32, Zero-shot=true2025.04 | 34.94 | 59.41 | 69.36 | |
| HyCD + LalignBackbone=ViT-B/32, Zero-shot=true2025.04 | 31.73 | 55.63 | 65.9 | |
| Pre-trained (CLIP)Backbone=ViT-B/32, Zero-shot=true2025.04 | 30.06 | 53.83 | 63.29 | |
| Self-KDBackbone=ViT-B/32, Zero-shot=true2025.04 | 29.98 | 53.83 | 63.55 |