Image-to-Text Retrieval on IIW (It Is What it is)
99.8R@1FineViT/14
Evaluation Results
| Method | Links | |
|---|---|---|
| FineViT/14Params=0.86B, Zero-shot=true2026.03 | 99.8 | |
| FixCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | 97.9 | |
| CAFT++Training Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 97.6 | |
| CAFTTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 97.4 | |
| SigLIP2-so400m/14Params=0.4B, Zero-shot=true2026.03 | 94.8 | |
| LongCLIP-L/14Params=0.3B, Zero-shot=true2026.03 | 94.1 | |
| LoTLIPTraining Data Scale=100M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 94 | |
| SigLIP-so400m/14Params=0.4B, Zero-shot=true2026.03 | 92.5 | |
| FLAIRTraining Data Scale=30M, Training Strategy=Trained on Long-Captions from Scratch, Vision Encoder=ViT-B/16, Zero-shot=true2026.02 | 91.3 |