Text-to-Image Retrieval on ShareGPT4V 10k
94.8Recall@1CAFT++
Evaluation Results
| Method | Links | |
|---|---|---|
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 94.8 | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 94.5 | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 89.4 | |
| LoTLIPTraining Data Scale=100M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 81.4 | |
| OpenCLIPTraining Data Scale=2B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 66.8 | |
| SigLIPTraining Data Scale=10B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 63 | |
| ALIGNTraining Data Scale=700M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 62.7 | |
| Long-CLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 62 | |
| LiTTraining Data Scale=100M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 50.6 |