Text-to-Image Retrieval on ShareGPT4V 1k
99Recall@1CAFT++
Evaluation Results
| Method | Links | |
|---|---|---|
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 99 | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 98.8 | |
| TULIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 98.6 | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 98 | |
| FG-CLIPTraining Data Scale=400M to 1.6B, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 94.9 | |
| OpenCLIPTraining Data Scale=2B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 87.7 | |
| Long-CLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 87.4 | |
| LoTLIPTraining Data Scale=100M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 86.8 | |
| ALIGNTraining Data Scale=700M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 85.3 | |
| SigLIPTraining Data Scale=10B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 83.4 | |
| LiTTraining Data Scale=100M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 80 |