Text-to-Image Retrieval on Urban-1K
98.8R@1LamRA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LamRAreranking=pointwise/listwise2024.12 | 98.8 | — | |
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 95.4 | — | |
| LamRA-Ret2024.12 | 95.1 | — | |
| HyFL-CLIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 94.3 | — | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 93.9 | — | |
| HiMo-CLIP†Backbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 93.2 | — | |
| FineLIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 93 | — | |
| β-CLIP (BCE)Data=WIT-400M + ShareGPT4V-1M, Objective=BCE2025.12 | 91.8 | — | |
| HyFL-CLIPBackbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 91.1 | — | |
| TULIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 91.1 | — | |
| LongD-CLIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 90.8 | — | |
| SmartCLIPBackbone=L/142025.07 | 90.1 | — | |
| SmartCLIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 90.1 | — | |
| FG-CLIP (w/ hard negatives)Data=WIT-400M + LAION-1.6B + GRIT-12M + 40M reg. + 10M hard negatives2025.12 | 89.9 | — | |
| FineLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 89.3 | — | |
| HiMo-CLIP*Backbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 89.2 | — | |
| β-CLIP (CE)Data=WIT-400M + ShareGPT4V-1M, Objective=CE2025.12 | 89 | — | |
| FLAIRData=30M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 87.7 | — | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 87.7 | — | |
| FLAIRData=DreamLIP 30M2025.12 | 87.7 | — | |
| Fix-CLIPBackbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 87.7 | — | |
| FineLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 87.4 | — | |
| Smart-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 87.4 | — | |
| SmartCLIPBackbone=B/162025.07 | 87.4 | — | |
| SmartCLIPBackbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 87.4 | — | |
| LongD-CLIPBackbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 87.3 | — | |
| FineLIP*Backbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 86.9 | — | |
| FLAIRData=15M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 86.6 | — | |
| TULIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 86.6 | — | |
| TULIPData=WIT-400M + ShareGPT4V-1M2025.12 | 86.6 | — | |
| TULIPBackbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 86.6 | — | |
| Long-CLIP†Backbone=ViT-L-14, Evaluation Protocol=Zero-shot2026.07 | 86.2 | — | |
| Long-CLIP-L2024.12 | 86.1 | — | |
| Long-CLIPBackbone=L/142025.07 | 86.1 | — | |
| E5-V2024.12 | 84 | — | |
| EVA-CLIPparameters=18B2024.12 | 81.7 | — | |
| Fix-CLIPBackbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 81.1 | — | |
| FLAIRData=12M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 80.6 | — | |
| Long-CLIP†Backbone=ViT-B-16, Evaluation Protocol=Zero-shot2026.07 | 79.6 | — | |
| Long-CLIPData=400M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 79.5 | — | |
| Long-CLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 79.5 | — | |
| Long-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 79.5 | — | |
| Long-CLIPBackbone=B/162025.07 | 79.5 | — | |
| EVA-CLIPparameters=8B2024.12 | 77.8 | — | |
| UniIR-CLIP2024.12 | 75 | — | |
| FLAIRData=3M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 69.5 | — | |
| OpenCLIPData=2B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 65.8 | — | |
| OpenCLIPTraining Data Scale=2B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 65.8 | — | |
| SigLIPData=10B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 62.1 | — | |
| SigLIPTraining Data Scale=10B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 62.1 | — | |
| SigLIPData=WebLI-1B2025.12 | 62.1 | — | |
| MagicLens-L2024.12 | 59.3 | — | |
| CLIPBackbone=B/162025.07 | 53.6 | — | |
| CLIPData=WIT-400M2025.12 | 53.2 | — | |
| CLIP-L2024.12 | 52.8 | — | |
| CLIPBackbone=L/142025.07 | 52.8 | — | |
| FineViT/14Params=0.86B, Zero-shot=true2026.03 | — | 99.1 | |
| FixCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | — | 96.3 | |
| LongCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | — | 84 | |
| OpenAI-CLIP-B/16re-caption=false2024.06 | — | 53.3 | |
| OpenCLIP-B/16re-caption=false2024.06 | — | 63.1 | |
| Recap-CLIP-B/16re-caption=false2024.06 | — | 50.9 | |
| Recap-CLIP-B/16re-caption=true2024.06 | — | 87.3 | |
| Recap-CLIP-L/14re-caption=false2024.06 | — | 64.6 | |
| Recap-CLIP-L/14re-caption=true2024.06 | — | 91.8 | |
| SigLIP-so400m/14Params=0.4B, Zero-shot=true2026.03 | — | 73.8 | |
| SigLIP2-so400m/14Params=0.4B, Zero-shot=true2026.03 | — | 75.6 |