Image-to-Text Retrieval on SV-1k
98.7R@1Smart-CLIP
Evaluation Results
| Method | Links | |
|---|---|---|
| Smart-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 98.7 | |
| TULIPData=WIT-400M + ShareGPT4V-1M2025.12 | 98.6 | |
| FLAIRData=30M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 98.5 | |
| FLAIRData=DreamLIP 30M2025.12 | 98.5 | |
| FLAIRData=15M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 97.4 | |
| FG-CLIP (w/ hard negatives)Data=WIT-400M + LAION-1.6B + GRIT-12M + 40M reg. + 10M hard negatives2025.12 | 96.7 | |
| FLAIRData=12M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 96.1 | |
| LoTLIPData=100M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 95.5 | |
| LoTLIPData=Mixed-100M2025.12 | 95.5 | |
| Long-CLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 94.7 | |
| β-CLIP (BCE)Data=WIT-400M + ShareGPT4V-1M, Objective=BCE2025.12 | 94.1 | |
| β-CLIP (CE)Data=WIT-400M + ShareGPT4V-1M, Objective=CE2025.12 | 94 | |
| FLAIRData=3M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 91 | |
| FineLIPData=WIT-400M + ShareGPT4V-1M2025.12 | 90.7 | |
| Long-CLIPData=400M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 90.6 | |
| EVA-CLIPData=Merged-2B2025.12 | 90.5 | |
| OpenCLIPData=2B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 90.3 | |
| ALIGNData=700M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 86.3 | |
| LiTData=100M, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 86 | |
| SigLIPData=10B, Zero-shot=true, Vision Encoder=ViT-B/162024.12 | 85.8 | |
| SigLIPData=WebLI-1B2025.12 | 85.8 | |
| CLIPData=WIT-400M2025.12 | 78.2 | |
| Fine-CLIPData=CC2.5M + 10.4M reg.2025.12 | 70.6 |