Text-to-Image Retrieval on Flickr30k (val)
77.6R@1FreqAdapter
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FreqAdapterFoundation Model=CLIP-L/14-3362026.05 | 77.6 | 95.06 | 97.44 | |
| CLIP-AdapterFoundation Model=CLIP-L/14-3362026.05 | 77.28 | 94.72 | 97.24 | |
| CLIP-AdapterFoundation Model=CLIP-L/142026.05 | 75.76 | 93.24 | 96.44 | |
| FreqAdapterFoundation Model=CLIP-L/142026.05 | 75.72 | 93.72 | 96.86 | |
| MMAFoundation Model=CLIP-L/14-3362026.05 | 75.68 | 93.66 | 96.6 | |
| CoOpFoundation Model=CLIP-L/14-3362026.05 | 75.62 | 94.22 | 97.28 | |
| CoOpFoundation Model=CLIP-L/142026.05 | 74.96 | 92.12 | 96.14 | |
| LoR-VPFoundation Model=CLIP-L/14-3362026.05 | 73.46 | 93.86 | 96.4 | |
| FreqAdapterFoundation Model=CLIP-B/162026.05 | 73.42 | 92.88 | 96.56 | |
| MMAFoundation Model=CLIP-L/142026.05 | 73.14 | 92.38 | 96.1 | |
| CoOpFoundation Model=CLIP-B/162026.05 | 72.04 | 93.22 | 96.26 | |
| CLIP-AdapterFoundation Model=CLIP-B/162026.05 | 71.26 | 92.64 | 96.08 | |
| MaPLeFoundation Model=CLIP-L/14-3362026.05 | 71.24 | 92.04 | 96.38 | |
| MaPLeFoundation Model=CLIP-L/142026.05 | 71.02 | 91.28 | 94.16 | |
| MMAFoundation Model=CLIP-B/162026.05 | 70.56 | 91.56 | 95.88 | |
| MaPLeFoundation Model=CLIP-B/162026.05 | 70.36 | 92.04 | 95.04 | |
| CLIP-L/14-336Foundation Model=CLIP-L/14-3362026.05 | 69.12 | 90.2 | 94.94 | |
| ITOTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 67.1 | 88.38 | 93.18 | |
| ITOPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 67.1 | 88.38 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 66.57 | 87.71 | 92.6 | |
| CLIPPre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 66.57 | 87.71 | — | |
| CLIP-L/14Foundation Model=CLIP-L/142026.05 | 66.36 | 88.5 | 93.62 | |
| ITOTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 66.33 | 87.69 | 92.96 | |
| ITOPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 66.33 | 87.69 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-B/16, Training Epochs=10, Zero-shot evaluation=true2026.03 | 65.88 | 87.67 | 92.49 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-B/16, Training Epochs=102026.03 | 65.88 | 87.67 | — | |
| ITOTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 63 | 86.21 | 91.42 | |
| ITOPre-training Scale=100M2026.03 | 63 | 86.21 | — | |
| ITO sub2Training Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 62.88 | 85.42 | 91.2 | |
| ITO sub2Pre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 62.88 | 85.42 | — | |
| CLIPTraining Dataset Size=1B, Backbone=ViT-L/16, Training Epochs=1, Zero-shot evaluation=true2026.03 | 62.6 | 85.62 | 91.26 | |
| CLIPPre-training Scale=1B, Backbone=ViT-L/16, Training Epochs=12026.03 | 62.6 | 85.62 | — | |
| CLIP-B/16Foundation Model=CLIP-B/162026.05 | 62.28 | 85.64 | 92.14 | |
| CLIPTraining Dataset Size=100M, Zero-shot evaluation=true2026.03 | 60.04 | 84.08 | 89.9 | |
| CLIPPre-training Scale=100M2026.03 | 60.04 | 84.08 | — | |
| ITO_sub3Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 59.13 | 83.41 | 89.8 | |
| ITO_sub2Pre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 57.32 | 82.23 | 88.38 | |
| ITOTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 56.79 | 82.17 | 88.56 | |
| ITOPre-training Scale=1B2026.03 | 56.79 | 82.17 | — | |
| ITO sub2Training Dataset Size=12M, Zero-shot evaluation=true2026.03 | 56.65 | 82.54 | 88.58 | |
| ITO sub2Pre-training Scale=12M2026.03 | 56.65 | 82.54 | — | |
| ITO sub2Training Dataset Size=1B, Zero-shot evaluation=true2026.03 | 55.7 | 82.33 | 89.33 | |
| ITO sub2Pre-training Scale=1B2026.03 | 55.7 | 82.33 | — | |
| CLIPTraining Dataset Size=1B, Zero-shot evaluation=true2026.03 | 54.73 | 80.49 | 87.53 | |
| CLIPPre-training Scale=1B2026.03 | 54.73 | 80.49 | — | |
| ITOTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 54.5 | 79.98 | 86.86 | |
| ITOPre-training Scale=12M2026.03 | 54.5 | 79.98 | — | |
| FLAIRPre-training dataset=CC3M-recap, Evaluation protocol=Zero-shot, Vision Encoder=ViT-B/162026.03 | 53.25 | 78.01 | 85.48 | |
| SigLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 50.87 | 75.96 | 84.18 | |
| SigLIPPre-training Scale=12M2026.03 | 50.87 | 75.96 | — | |
| FLAIRTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 47.36 | 74.77 | 83.21 | |
| FLAIRPre-training Scale=12M2026.03 | 47.36 | 74.77 | — | |
| CLIPTraining Dataset Size=12M, Zero-shot evaluation=true2026.03 | 45.74 | 73.02 | 82.23 | |
| CLIPPre-training Scale=12M2026.03 | 45.74 | 73.02 | — | |
| ITO sub2Training Dataset Size=15M, Zero-shot evaluation=true2026.03 | 37.95 | 65.46 | 75.5 | |
| ITO sub2Pre-training Scale=15M2026.03 | 37.95 | 65.46 | — | |
| ITOTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 35.72 | 62.6 | 73.06 | |
| ITOPre-training Scale=15M2026.03 | 35.72 | 62.6 | — | |
| CLIPTraining Dataset Size=15M, Zero-shot evaluation=true2026.03 | 30.77 | 56.82 | 67.14 | |
| CLIPPre-training Scale=15M2026.03 | 30.77 | 56.82 | — | |
| ITOTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 29.45 | 53.89 | 64.36 | |
| ITOPre-training Scale=3M2026.03 | 29.45 | 53.89 | — | |
| ITO sub2Training Dataset Size=3M, Zero-shot evaluation=true2026.03 | 28.26 | 53.08 | 64.16 | |
| ITO sub2Pre-training Scale=3M2026.03 | 28.26 | 53.08 | — | |
| FLAIRTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 25.74 | 49.86 | 60.81 | |
| FLAIRPre-training Scale=3M2026.03 | 25.74 | 49.86 | — | |
| SigLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 18.32 | 39.43 | 49.45 | |
| SigLIPPre-training Scale=3M2026.03 | 18.32 | 39.43 | — | |
| CLIPTraining Dataset Size=3M, Zero-shot evaluation=true2026.03 | 16.98 | 37.65 | 48.68 | |
| CLIPPre-training Scale=3M2026.03 | 16.98 | 37.65 | — |