Image-to-Text Retrieval on Urban-1K
98R@1LamRA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LamRAreranking=pointwise/listwise2024.12 | 98 | — | |
| LamRA-Ret2024.12 | 94.3 | — | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 93.6 | — | |
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 93.6 | — | |
| FG-CLIP (w/ hard negatives)Data=WIT-400M + LAION-1.6B + GRIT-12M + 40M reg. + 10M hard negatives2025.12 | 93 | — | |
| β-CLIP (BCE)Data=WIT-400M + ShareGPT4V-1M, Objective=BCE2025.12 | 92.3 | — | |
| FineLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 90.7 | — | |
| FineLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 90 | — | |
| Smart-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 90 | — | |
| β-CLIP (CE)Data=WIT-400M + ShareGPT4V-1M, Objective=CE2025.12 | 88.6 | — | |
| TULIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 88.1 | — | |
| TULIPData=WIT-400M + ShareGPT4V-1M2025.12 | 88.1 | — | |
| FLAIRData=30M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 83.6 | — | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 83.6 | — | |
| FLAIRData=DreamLIP 30M2025.12 | 83.6 | — | |
| EVA-CLIPparameters=18B2024.12 | 83.3 | — | |
| Long-CLIP-L2024.12 | 82.7 | — | |
| FLAIRData=15M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 82.4 | — | |
| E5-V2024.12 | 82.4 | — | |
| EVA-CLIPparameters=8B2024.12 | 80.4 | — | |
| Long-CLIPData=400M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 78.9 | — | |
| Long-CLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 78.9 | — | |
| Long-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 78.9 | — | |
| UniIR-CLIP2024.12 | 78.4 | — | |
| FLAIRData=12M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 74.6 | — | |
| OpenCLIPData=2B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 69.5 | — | |
| OpenCLIPTraining Data Scale=2B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 69.5 | — | |
| CLIP-L2024.12 | 68.7 | — | |
| CLIPData=WIT-400M2025.12 | 67.5 | — | |
| FLAIRData=3M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 63.5 | — | |
| SigLIPData=10B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 62.7 | — | |
| SigLIPTraining Data Scale=10B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 62.7 | — | |
| SigLIPData=WebLI-1B2025.12 | 62.7 | — | |
| MagicLens-L2024.12 | 24.2 | — | |
| FineViT/14Params=0.86B, Zero-shot=true2026.03 | — | 98.9 | |
| FixCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | — | 93.7 | |
| LongCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | — | 81.1 | |
| OpenAI-CLIP-B/16re-caption=false2024.06 | — | 67.4 | |
| OpenCLIP-B/16re-caption=false2024.06 | — | 62.5 | |
| Recap-CLIP-B/16re-caption=false2024.06 | — | 53.2 | |
| Recap-CLIP-B/16re-caption=true2024.06 | — | 85 | |
| Recap-CLIP-L/14re-caption=false2024.06 | — | 69.8 | |
| Recap-CLIP-L/14re-caption=true2024.06 | — | 89 | |
| SigLIP-so400m/14Params=0.4B, Zero-shot=true2026.03 | — | 74.7 | |
| SigLIP2-so400m/14Params=0.4B, Zero-shot=true2026.03 | — | 78.1 |