Image-to-Text Retrieval on ShareGPT4V 1k
99.5R@1CAFT++
Evaluation Results
| Method | Links | |
|---|---|---|
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 99.5 | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 99 | |
| TULIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 98.6 | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 98.5 | |
| FG-CLIPTraining Data Scale=400M to 1.6B, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 96.7 | |
| LoTLIPTraining Data Scale=100M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 95.5 | |
| Long-CLIPTraining Data Scale=400M to 1M, Training Strategy=Finetuned on Long-Captions, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 90.6 | |
| OpenCLIPTraining Data Scale=2B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 90.3 | |
| ALIGNTraining Data Scale=700M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 86.3 | |
| LiTTraining Data Scale=100M, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 86 | |
| SigLIPTraining Data Scale=10B, Training Strategy=Trained on Short-Captions Only, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 85.8 |