Text-to-Image Retrieval on MSCOCO (val)
38.97R@1ITO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ITOTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 38.97 | 64.45 | 74.63 | |
| ITOPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 38.97 | 64.45 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 38.33 | 64.03 | 74.03 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 38.33 | 64.03 | — | |
| ITOTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 38.12 | 63.72 | 73.62 | |
| ITOPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 38.12 | 63.72 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 37.76 | 63.58 | 73.5 | |
| CLIPPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 37.76 | 63.58 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 35.69 | 61.62 | 71.86 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 35.69 | 61.62 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 34.88 | 60.16 | 71.04 | |
| CLIPPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 34.88 | 60.16 | — | |
| ITOTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 34.26 | 59.57 | 70.07 | |
| ITOPre-training Scale=100M2026.03 | 34.26 | 59.57 | — | |
| ITO_sub3Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 32.38 | 58.74 | 69.77 | |
| CLIPTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 31.92 | 57.32 | 68.4 | |
| CLIPPre-training Scale=100M2026.03 | 31.92 | 57.32 | — | |
| ITO sub2Training Dataset Size=1B, Zero-shot evaluation=true2026.03 | 31.89 | 57.26 | 68.36 | |
| ITO sub2Pre-training Scale=1B2026.03 | 31.89 | 57.26 | — | |
| ITO_sub2Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 31.59 | 58.54 | 69.98 | |
| ITOTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 31.3 | 56.63 | 68.22 | |
| ITOPre-training Scale=1B2026.03 | 31.3 | 56.63 | — | |
| ITO sub2Training Dataset Size=12M, Zero-shot evaluation=true2026.03 | 30.51 | 56.76 | 68.29 | |
| ITO sub2Pre-training Scale=12M2026.03 | 30.51 | 56.76 | — | |
| FLAIRPre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 29.76 | 55.53 | 66.69 | |
| ITOTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 29.62 | 55.74 | 67.09 | |
| ITOPre-training Scale=12M2026.03 | 29.62 | 55.74 | — | |
| CLIPTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 29.51 | 54.84 | 66.05 | |
| CLIPPre-training Scale=1B2026.03 | 29.51 | 54.84 | — | |
| SigLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 26.85 | 51.48 | 63.2 | |
| SigLIPPre-training Scale=12M2026.03 | 26.85 | 51.48 | — | |
| FLAIRTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 24.62 | 48.79 | 60.57 | |
| FLAIRPre-training Scale=12M2026.03 | 24.62 | 48.79 | — | |
| CLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 23.04 | 47.58 | 59.39 | |
| CLIPPre-training Scale=12M2026.03 | 23.04 | 47.58 | — | |
| ITO sub2Training Dataset Size=15M, Zero-shot evaluation=true2026.03 | 20.04 | 42.88 | 54.61 | |
| ITO sub2Pre-training Scale=15M2026.03 | 20.04 | 42.88 | — | |
| ITOTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 19.57 | 41.61 | 53.03 | |
| ITOPre-training Scale=15M2026.03 | 19.57 | 41.61 | — | |
| CLIPTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 15.06 | 34.82 | 46.43 | |
| CLIPPre-training Scale=15M2026.03 | 15.06 | 34.82 | — | |
| ITOTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 14.56 | 33.71 | 44.9 | |
| ITOPre-training Scale=3M2026.03 | 14.56 | 33.71 | — | |
| ITO sub2Training Dataset Size=3M, Zero-shot evaluation=true2026.03 | 14.01 | 33.88 | 45.06 | |
| ITO sub2Pre-training Scale=3M2026.03 | 14.01 | 33.88 | — | |
| FLAIRTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 12.59 | 30.38 | 40.94 | |
| FLAIRPre-training Scale=3M2026.03 | 12.59 | 30.38 | — | |
| SigLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 9.28 | 23.48 | 32.6 | |
| SigLIPPre-training Scale=3M2026.03 | 9.28 | 23.48 | — | |
| CLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 8.14 | 22.68 | 31.98 | |
| CLIPPre-training Scale=3M2026.03 | 8.14 | 22.68 | — |