Image Retrieval on MS-COCO 5K (test)
67.2R@1CoCa
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CoCaParameters=2.1B, Evaluation Protocol=Fine-Tuned, Training Data=Super-Large2023.01 | 67.2 | 87.7 | 92.8 | |
| X-FMbaseParameters=327M, Evaluation Protocol=Fine-Tuned, Training Data=More Data2023.01 | 67 | 87.2 | 92.4 | |
| X2-VLMbaseParameters=255M, Evaluation Protocol=Fine-Tuned, Training Data=More Data2023.01 | 66.2 | 87.1 | 92.2 | |
| mPLUG-2#PT Data=17M2023.02 | 65.7 | 87.1 | 92.6 | |
| mPLUG-2Base#PT Data=17M2023.02 | 65.3 | 86.9 | 92.4 | |
| VLMO-Large++Pretrain Images=1.0B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product), large batch size=true2021.11 | 65.2 | 86.5 | 92.2 | |
| BLIP#PT Data=129M2023.02 | 65.1 | 86.3 | 91.8 | |
| BLIP2023.05 | 65.1 | 86.3 | 91.8 | |
| OmniVL# Img-Text Pairs=14M*, Fine-tuned=true2022.09 | 64.8 | 86.1 | 91.6 | |
| X-FMbaseParameters=327M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 64.7 | 86.1 | 91.6 | |
| BLIPPre-training data size=129M2022.06 | 64.3 | 85.7 | 91.5 | |
| X-VLM + EPICData=4M+, EPIC=true, Protocol=Fine-tuned2022.11 | 64.1 | 86.1 | 91.6 | |
| X-VLM + EPICData=16M+, EPIC=true, Protocol=Fine-tuned2022.11 | 64.1 | 85.9 | 91.8 | |
| X-VLM# Params=216M, # Pre-train Images=16M2021.11 | 63.4 | 85.8 | 91.5 | |
| X-VLMData=16M+, EPIC=false, Protocol=Fine-tuned2022.11 | 63.3 | 85.6 | 91.4 | |
| Florence-HugePretrain Images=900M, Model Size Category=Huge-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 63.2 | 85.7 | — | |
| Florence#PT Data=0.9B2023.02 | 63.2 | 85.7 | — | |
| Florence# Img-Text Pairs=900M, Fine-tuned=true2022.09 | 63.2 | 85.7 | — | |
| BLIPPre-training data size=14M2022.06 | 63.1 | 85.3 | 91.1 | |
| X-VLM# Params=216M, # Pre-train Images=4M2021.11 | 63.1 | 85.7 | 91.6 | |
| X-VLMParameters=216M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 63.1 | 85.7 | 91.6 | |
| BLIP# Img-Text Pairs=14M, Fine-tuned=true2022.09 | 63.1 | 85.3 | 91.1 | |
| ALBEF + EPICData=16M, EPIC=true, Protocol=Fine-tuned2022.11 | 62.9 | 85.4 | 91.3 | |
| X-VLMData=4M+, EPIC=false, Protocol=Fine-tuned2022.11 | 62.7 | 85.6 | 91.4 | |
| X2-VLMbaseParameters=255M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 62.7 | 84.7 | 90.7 | |
| METER + EPICData=16M, EPIC=true, Protocol=Fine-tuned2022.11 | 62.5 | 85.4 | 91.9 | |
| VL-BEITParameters=175M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 61.5 | — | — | |
| ALBEFData=16M, EPIC=false, Protocol=Fine-tuned2022.11 | 61.3 | 84.3 | 90.6 | |
| METER + EPICData=4M, EPIC=true, Protocol=Fine-tuned2022.11 | 61.2 | 85.2 | 91.6 | |
| X-FMbaseParameters=327M, Evaluation Protocol=Zero-Shot, Training Data=More Data2023.01 | 61.1 | 84.5 | 90.6 | |
| BEIT-3Parameters=1.9B, Evaluation Protocol=Zero-Shot, Training Data=Super-Large2023.01 | 61.1 | 84.5 | 90.6 | |
| RaSa2023.05 | 61 | 84.49 | 90.83 | |
| CCLM_basePre-training Data=4M2022.06 | 60.89 | — | — | |
| ALBEFPre-train Images=14M, Fine-tuned=true2021.07 | 60.7 | 84.3 | 90.5 | |
| ALBEFPre-training data size=14M2022.06 | 60.7 | 84.3 | 90.5 | |
| ALBEF# Params=210M, # Pre-train Images=14M2021.11 | 60.7 | 84.3 | 90.5 | |
| ALBEF#PT Data=14M2023.02 | 60.7 | 84.3 | 90.5 | |
| ALBEF# Img-Text Pairs=14M, Fine-tuned=true2022.09 | 60.7 | 84.3 | 90.5 | |
| VLMO-LargePretrain Images=4M, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 60.6 | 84.4 | 91 | |
| VLMo#PT Data=4M2023.02 | 60.6 | 84.4 | 91 | |
| AlignCMSS2023.09 | 60.4 | 84.3 | 90.7 | |
| ALBEFconfiguration=backbone2023.05 | 60.31 | 84.22 | 90.51 | |
| ALIGNPre-train Images=1.2B, Fine-tuned=true2021.07 | 59.9 | 83.3 | 89.8 | |
| ALIGN#Images=1.2B, Fine-tuning=true2022.02 | 59.9 | 83.3 | 89.8 | |
| ALIGN-LargePretrain Images=1.8B, Model Size Category=Large-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 59.9 | 83.3 | 89.8 | |
| ALIGNPre-training data size=1.2B2022.06 | 59.9 | 83.3 | 89.8 | |
| ALIGN# Params=490M, # Pre-train Images=1.8B2021.11 | 59.9 | 83.3 | 89.8 | |
| ALIGNTraining Data=1.8B2022.07 | 59.9 | 83.3 | 89.8 | |
| ALIGN#PT Data=1.8B2023.02 | 59.9 | 83.3 | 89.8 | |
| ALIGN2023.05 | 59.9 | 83.3 | 89.8 | |
| ALIGN# Img-Text Pairs=1.8B, Fine-tuned=true2022.09 | 59.9 | 83.3 | 89.8 | |
| METERData=16M, EPIC=false, Protocol=Fine-tuned2022.11 | 59.8 | 84.2 | 90.8 | |
| SINGULARITYPre-training data size=17M2022.06 | 59.6 | 83.4 | 90 | |
| X-FMbaseParameters=327M, Evaluation Protocol=Zero-Shot, Training Data=Standard2023.01 | 59.4 | 83.6 | 90 | |
| METERData=4M, EPIC=false, Protocol=Fine-tuned2022.11 | 59.2 | 84 | 90.8 | |
| Triple Contrastive Learning (TCL)#Images=4M, Fine-tuning=true2022.02 | 59 | 83.2 | 89.9 | |
| OSCAR+ w/ VINVLBERT scale=Large2021.01 | 58.8 | 83.5 | 90.3 | |
| VinVL-LargePretrain Images=5.7M, Model Size Category=Large-Size, Interaction Protocol=Fusion encoder2021.11 | 58.8 | 83.5 | 90.3 | |
| VinVL_large# Params=550M, # Pre-train Images=5.6M2021.11 | 58.8 | 83.5 | 90.3 | |
| ALBEF + EPICData=4M, EPIC=true, Protocol=Fine-tuned2022.11 | 58.6 | 82.7 | 89.3 | |
| OmniVLParameters=288M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 58.5 | 82.6 | 89.5 | |
| OmniVL# Img-Text Pairs=4M*, Fine-tuned=true2022.09 | 58.5 | 82.6 | 89.5 | |
| X2-VLMbaseParameters=255M, Evaluation Protocol=Zero-Shot, Training Data=More Data2023.01 | 58.3 | 84.7 | 91 | |
| OSCAR+ w/ VINVLBERT scale=Base2021.01 | 58.1 | 83.2 | 90.1 | |
| VinVL-BaseVisual Embed=Region, Time (ms)=~650, Extra pre-training=GQA, VQAv2, VG-QA, Extra pre-training dataset=Open Images2021.02 | 58.1 | 83.2 | 90.1 | |
| VinVL (Base)Training Data=8.9M2022.07 | 58.1 | 83.2 | 90.1 | |
| VinVL_baseModel Scale=base2022.06 | 58.1 | — | — | |
| VinVL2023.09 | 58.1 | 83.2 | 90.1 | |
| Knowledge-CLIPMode=Fine-tuned2022.10 | 57.6 | 83.9 | 90.4 | |
| OSCARModel Size=Large2020.04 | 57.5 | 82.8 | 89.8 | |
| OSCARSize=L2020.04 | 57.5 | 82.8 | 89.8 | |
| OscarEvaluation Protocol=Finetune2021.02 | 57.5 | 82.8 | 89.8 | |
| OscarBERT scale=Large2021.01 | 57.5 | 82.8 | 89.8 | |
| OSCARMode=Fine-tuned2022.10 | 57.5 | 82.8 | 89.8 | |
| X-VLMData=16M+, EPIC=true, Evaluation Protocol=zero-shot2022.11 | 57.5 | 83.3 | 90 | |
| X-VLM + EPICData=16M+, EPIC=true, Protocol=Zero-shot2022.11 | 57.5 | 83.3 | 90 | |
| Oscar2023.05 | 57.5 | 82.8 | 89.8 | |
| X-VLMData=4M+, EPIC=true, Evaluation Protocol=zero-shot2022.11 | 57.3 | 83.4 | 90.2 | |
| X-VLM + EPICData=4M+, EPIC=true, Protocol=Zero-shot2022.11 | 57.3 | 83.4 | 90.2 | |
| VLMO-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Dual encoder (dot product)2021.11 | 57.2 | 82.6 | 89.8 | |
| VLMbaseParameters=175M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 57.2 | 82.6 | 89.8 | |
| VLMO# Img-Text Pairs=4M, Fine-tuned=true2022.09 | 57.2 | 82.6 | 89.8 | |
| METER-CLIP# Params=380M, # Pre-train Images=4M2021.11 | 57.1 | 82.7 | 90.1 | |
| METER# Img-Text Pairs=404M, Fine-tuned=true2022.09 | 57.1 | 82.7 | 90.1 | |
| X-VLMData=16M+, EPIC=false, Evaluation Protocol=zero-shot2022.11 | 56.9 | 83 | 89.9 | |
| X-VLMData=16M+, EPIC=false, Protocol=Zero-shot2022.11 | 56.9 | 83 | 89.9 | |
| ALBEFPre-train Images=4M, Fine-tuned=true2021.07 | 56.8 | 81.5 | 89.2 | |
| ALBEF#Images=4M, Fine-tuning=true2022.02 | 56.8 | 81.5 | 89.2 | |
| ALBEF-BasePretrain Images=4M, Model Size Category=Base-Size, Interaction Protocol=Fusion reranking2021.11 | 56.8 | 81.5 | 89.2 | |
| ALBEFPre-training data size=4M2022.06 | 56.8 | 81.5 | 89.2 | |
| ALBEF# Params=210M, # Pre-train Images=4M2021.11 | 56.8 | 81.5 | 89.2 | |
| ALBEFPre-training Data=4M2022.06 | 56.8 | — | — | |
| ALBEFParameters=210M, Evaluation Protocol=Fine-Tuned, Training Data=Standard2023.01 | 56.8 | 81.5 | 89.2 | |
| ALBEFData=4M, EPIC=false, Protocol=Fine-tuned2022.11 | 56.4 | 81.8 | 88.9 | |
| X-VLM# Params=216M, # Pre-train Images=16M, Zero-shot=true2021.11 | 56.1 | 83 | 89.8 | |
| X-VLM# Params=216M, # Pre-train Images=4M, Zero-shot=true2021.11 | 55.6 | 82.7 | 90 | |
| X-VLMParameters=216M, Evaluation Protocol=Zero-Shot, Training Data=Standard2023.01 | 55.6 | 82.7 | 90 | |
| X-VLMData=4M+, EPIC=false, Evaluation Protocol=zero-shot2022.11 | 55.3 | 82.5 | 89.7 | |
| X-VLMData=4M+, EPIC=false, Protocol=Zero-shot2022.11 | 55.3 | 82.5 | 89.7 | |
| X2-VLMbaseParameters=255M, Evaluation Protocol=Zero-Shot, Training Data=Standard2023.01 | 55.2 | 82.2 | 89.3 |