Image-to-Text Retrieval on MSCOCO (val)
58.14R@1ITO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ITOTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 58.14 | 81.76 | 88.88 | |
| ITOPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 58.14 | 81.76 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 57.7 | 80.86 | 87.72 | |
| CLIPPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 57.7 | 80.86 | — | |
| ITOTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 56.58 | 80.46 | 88.12 | |
| ITOPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 56.58 | 80.46 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 55.88 | 79.9 | 87.84 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 55.88 | 79.9 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 53.4 | 77.76 | 86.62 | |
| CLIPPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 53.4 | 77.76 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 52.12 | 77.44 | 85.96 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 52.12 | 77.44 | — | |
| ITOTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 52.08 | 76.04 | 84.24 | |
| ITOPre-training Scale=100M2026.03 | 52.08 | 76.04 | — | |
| ITO sub2Training Dataset Size=1B, Zero-shot evaluation=true2026.03 | 49.5 | 74.12 | 83.06 | |
| ITO sub2Pre-training Scale=1B2026.03 | 49.5 | 74.12 | — | |
| ITOTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 49.26 | 74.76 | 83.62 | |
| ITOPre-training Scale=1B2026.03 | 49.26 | 74.76 | — | |
| CLIPTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 49.12 | 74.66 | 83.76 | |
| CLIPPre-training Scale=100M2026.03 | 49.12 | 74.66 | — | |
| ITO_sub2Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 47.76 | 73.76 | 82.64 | |
| ITO_sub3Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 47.4 | 74.94 | 83.86 | |
| CLIPTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 47.08 | 72.7 | 82.64 | |
| CLIPPre-training Scale=1B2026.03 | 47.08 | 72.7 | — | |
| ITO sub2Training Dataset Size=12M, Zero-shot evaluation=true2026.03 | 43.94 | 70.4 | 80.26 | |
| ITO sub2Pre-training Scale=12M2026.03 | 43.94 | 70.4 | — | |
| ITOTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 42.3 | 70 | 79.46 | |
| ITOPre-training Scale=12M2026.03 | 42.3 | 70 | — | |
| SigLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 39.98 | 67.28 | 77.82 | |
| SigLIPPre-training Scale=12M2026.03 | 39.98 | 67.28 | — | |
| FLAIRPre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 37.54 | 64.38 | 75.88 | |
| FLAIRTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 36.2 | 63.16 | 74.38 | |
| FLAIRPre-training Scale=12M2026.03 | 36.2 | 63.16 | — | |
| CLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 34.17 | 61.2 | 72.52 | |
| CLIPPre-training Scale=12M2026.03 | 34.17 | 61.2 | — | |
| ITO sub2Training Dataset Size=15M, Zero-shot evaluation=true2026.03 | 31.42 | 58.02 | 69.06 | |
| ITO sub2Pre-training Scale=15M2026.03 | 31.42 | 58.02 | — | |
| ITOTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 30.76 | 57.44 | 68.82 | |
| ITOPre-training Scale=15M2026.03 | 30.76 | 57.44 | — | |
| CLIPTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 26.38 | 50.86 | 62.86 | |
| CLIPPre-training Scale=15M2026.03 | 26.38 | 50.86 | — | |
| ITOTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 21.56 | 45.36 | 57.22 | |
| ITOPre-training Scale=3M2026.03 | 21.56 | 45.36 | — | |
| ITO sub2Training Dataset Size=3M, Zero-shot evaluation=true2026.03 | 21.08 | 45.62 | 57.42 | |
| ITO sub2Pre-training Scale=3M2026.03 | 21.08 | 45.62 | — | |
| FLAIRTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 17.86 | 38.9 | 51.18 | |
| FLAIRPre-training Scale=3M2026.03 | 17.86 | 38.9 | — | |
| SigLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 13.7 | 32.66 | 43.16 | |
| SigLIPPre-training Scale=3M2026.03 | 13.7 | 32.66 | — | |
| CLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 12.36 | 30.98 | 41.76 | |
| CLIPPre-training Scale=3M2026.03 | 12.36 | 30.98 | — |