Image-to-Text Retrieval on Flickr30k (val)
90.9Recall@1FreqAdapter
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FreqAdapterFoundation Model=CLIP-L/14-3362026.05 | 90.9 | 98.7 | 99.7 | |
| LoR-VPFoundation Model=CLIP-L/14-3362026.05 | 90.3 | 98.3 | 99.5 | |
| CoOpFoundation Model=CLIP-L/14-3362026.05 | 90.2 | 98.7 | 99.6 | |
| CLIP-AdapterFoundation Model=CLIP-L/14-3362026.05 | 90 | 98.4 | 99.6 | |
| MMAFoundation Model=CLIP-L/14-3362026.05 | 89.9 | 99.1 | 99.8 | |
| CLIP-L/14-336Foundation Model=CLIP-L/14-3362026.05 | 89.8 | 99.1 | 99.8 | |
| MaPLeFoundation Model=CLIP-L/14-3362026.05 | 88.3 | 97.5 | 99.1 | |
| CoOpFoundation Model=CLIP-L/142026.05 | 87.7 | 98.5 | 99.3 | |
| FreqAdapterFoundation Model=CLIP-L/142026.05 | 87.5 | 98.7 | 99.6 | |
| CLIP-AdapterFoundation Model=CLIP-L/142026.05 | 87.3 | 98.2 | 99.3 | |
| MMAFoundation Model=CLIP-L/142026.05 | 87.2 | 98.1 | 99.3 | |
| FreqAdapterFoundation Model=CLIP-B/162026.05 | 86.8 | 98.5 | 99.3 | |
| CLIP-L/14Foundation Model=CLIP-L/142026.05 | 86.6 | 97.9 | 99.4 | |
| MaPLeFoundation Model=CLIP-L/142026.05 | 86.1 | 98 | 99.2 | |
| MMAFoundation Model=CLIP-B/162026.05 | 86 | 97.8 | 98.9 | |
| CLIP-B/16Foundation Model=CLIP-B/162026.05 | 85.3 | 97 | 98.6 | |
| ITOTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 84.81 | 96.45 | 98.52 | |
| ITOPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 84.81 | 96.45 | — | |
| CoOpFoundation Model=CLIP-B/162026.05 | 84.6 | 97.5 | 99.1 | |
| CLIP-AdapterFoundation Model=CLIP-B/162026.05 | 83.9 | 97.6 | 99.2 | |
| ITOTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 83.14 | 96.35 | 98.62 | |
| ITOPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 83.14 | 96.35 | — | |
| MaPLeFoundation Model=CLIP-B/162026.05 | 82.4 | 97.2 | 99 | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 82.25 | 96.65 | 98.52 | |
| CLIPPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 82.25 | 96.65 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 81.95 | 95.36 | 98.03 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 81.95 | 95.36 | — | |
| ITOTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 79.59 | 95.66 | 98.42 | |
| ITOPre-training Scale=100M2026.03 | 79.59 | 95.66 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 78.99 | 95.27 | 98.03 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 78.99 | 95.27 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 78.4 | 95.07 | 97.63 | |
| CLIPPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 78.4 | 95.07 | — | |
| CLIPTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 75.94 | 93.59 | 96.75 | |
| ITOTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 75.94 | 94.08 | 97.14 | |
| CLIPPre-training Scale=100M2026.03 | 75.94 | 93.59 | — | |
| ITOPre-training Scale=1B2026.03 | 75.94 | 94.08 | — | |
| ITO_sub3Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 74.95 | 94.38 | 96.75 | |
| ITO_sub2Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 74.56 | 93.89 | 96.75 | |
| ITO sub2Training Dataset Size=1B, Zero-shot evaluation=true2026.03 | 73.37 | 93 | 96.15 | |
| ITO sub2Pre-training Scale=1B2026.03 | 73.37 | 93 | — | |
| CLIPTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 72.68 | 90.63 | 94.48 | |
| CLIPPre-training Scale=1B2026.03 | 72.68 | 90.63 | — | |
| ITOTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 72.49 | 90.83 | 95.07 | |
| ITOPre-training Scale=12M2026.03 | 72.49 | 90.83 | — | |
| ITO sub2Training Dataset Size=12M, Zero-shot evaluation=true2026.03 | 72.29 | 92.8 | 96.35 | |
| ITO sub2Pre-training Scale=12M2026.03 | 72.29 | 92.8 | — | |
| SigLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 68.34 | 89.45 | 93.29 | |
| SigLIPPre-training Scale=12M2026.03 | 68.34 | 89.45 | — | |
| FLAIRPre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 65.19 | 87.08 | 92.41 | |
| FLAIRTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 62.92 | 88.56 | 93.2 | |
| FLAIRPre-training Scale=12M2026.03 | 62.92 | 88.56 | — | |
| CLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 62.23 | 86.29 | 92.11 | |
| CLIPPre-training Scale=12M2026.03 | 62.23 | 86.29 | — | |
| ITO sub2Training Dataset Size=15M, Zero-shot evaluation=true2026.03 | 56.21 | 83.14 | 90.43 | |
| ITO sub2Pre-training Scale=15M2026.03 | 56.21 | 83.14 | — | |
| ITOTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 54.83 | 82.74 | 88.56 | |
| ITOPre-training Scale=15M2026.03 | 54.83 | 82.74 | — | |
| CLIPTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 47.14 | 74.46 | 83.23 | |
| CLIPPre-training Scale=15M2026.03 | 47.14 | 74.46 | — | |
| ITOTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 42.6 | 71.2 | 80.87 | |
| ITOPre-training Scale=3M2026.03 | 42.6 | 71.2 | — | |
| ITO sub2Training Dataset Size=3M, Zero-shot evaluation=true2026.03 | 42.11 | 69.43 | 78.11 | |
| ITO sub2Pre-training Scale=3M2026.03 | 42.11 | 69.43 | — | |
| FLAIRTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 35.4 | 62.13 | 74.46 | |
| FLAIRPre-training Scale=3M2026.03 | 35.4 | 62.13 | — | |
| SigLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 28.9 | 53.25 | 65.09 | |
| SigLIPPre-training Scale=3M2026.03 | 28.9 | 53.25 | — | |
| CLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 23.57 | 50.69 | 62.03 | |
| CLIPPre-training Scale=3M2026.03 | 23.57 | 50.69 | — |