Text Retrieval on MS-COCO
96.4R@5X-FM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| X-FMEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | 96.4 | 84.2 | 98.4 | |
| X2-VLMEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | 96.3 | 83.5 | 98.5 | |
| X-VLMEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | 95.5 | 80.4 | 98.2 | |
| mPLUG-2Evaluation protocol=Fine-Tune, Model size=base-size2023.01 | 95.2 | 81.2 | 98.1 | |
| X-FMEvaluation protocol=Zero-Shot, Model size=base-size2023.01 | 94.8 | 77.6 | 97.7 | |
| OmniVLEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | 93.6 | 76.8 | 97.3 | |
| OSCARLModel scale=Large2020.04 | 92.2 | 73.5 | 96 | |
| X-VLMEvaluation protocol=Zero-Shot, Model size=base-size2023.01 | 92.1 | 70.8 | 96.5 | |
| OSCARBModel scale=Base2020.04 | 91.1 | 70 | 95.5 | |
| SoTALModel scale=Large2020.04 | 89.4 | 66.6 | 94.3 | |
| SoTABModel scale=Base2020.04 | 87 | 63.3 | 93.1 | |
| SoTASModel scale=Small2020.04 | 84.5 | 56.6 | 92 | |
| FLAVAEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | 82.1 | 61.5 | 89.6 | |
| BaselineEvaluation protocol=finetuned2024.11 | 81.4 | — | — | |
| DCRZero-shot=true2026.03 | 79.3 | — | — | |
| Original CLIPZero-shot=true2026.03 | 79.2 | — | — | |
| FLAVAEvaluation protocol=Zero-Shot, Model size=base-size2023.01 | 76.8 | 42.7 | — | |
| CLIPEvaluation protocol=standard2024.11 | 74.9 | — | — | |
| TSVLC (LLM+RB)Method variant=LLM+RB2024.11 | 71.82 | — | — | |
| TSVLC (RB)Method variant=RB2024.11 | 71.7 | — | — | |
| CLIPEvaluation protocol=finetuned2024.11 | 68.9 | — | — | |
| NegCLIP2024.11 | 66 | — | — | |
| CLIP-PGSImage Enc.=ViT-B/16, Ratio=0.3, Zero-shot=true2025.03 | 64.4 | 36 | 74.6 | |
| CLIPImage Enc.=ViT-B/16, Zero-shot=true2025.03 | 62 | 34.6 | 72.7 | |
| E-CLIPImage Enc.=ViT-B/16, Zero-shot=true2025.03 | 62 | 34.3 | 73.3 | |
| CLIP-PGSImage Enc.=ViT-B/16, Ratio=0.5, Zero-shot=true2025.03 | 61.9 | 35.2 | 72.8 | |
| GenHancerZero-shot=true2026.03 | 61.4 | — | — | |
| CLIP-PGSImage Enc.=ViT-S/16, Ratio=0.3, Zero-shot=true2025.03 | 60.5 | 33.1 | 72.4 | |
| A-CLIPImage Enc.=ViT-B/16, Zero-shot=true2025.03 | 60.2 | 33.7 | 71 | |
| FLIPImage Enc.=ViT-B/16, Zero-shot=true2025.03 | 59.1 | 32.6 | 70.6 | |
| CLIP-PGSImage Enc.=ViT-S/16, Ratio=0.5, Zero-shot=true2025.03 | 58.1 | 31.6 | 69.6 | |
| TripletCLIPEvaluation protocol=finetuned2024.11 | 55.6 | — | — | |
| DAC2024.11 | 54.5 | — | — | |
| CyCLIPsmode=zero-shot2023.10 | 44.34 | 21.3 | 56.54 | |
| CyCLIPmode=zero-shot2023.10 | 41.46 | 18.92 | 54 | |
| CLIPsmode=zero-shot2023.10 | 38.92 | 17.78 | 50.1 | |
| CLIPmode=zero-shot2023.10 | 37.22 | 15.7 | 49.06 | |
| CyCLIPnmode=zero-shot2023.10 | 36.76 | 16.32 | 48.16 | |
| CLIPnmode=zero-shot2023.10 | 35.66 | 15.74 | 47.38 | |
| Flamingo-3Bparameters=3.2B, task-specific fine-tuning=false2022.11 | — | 65.9 | — | |
| Uni-Perceiver BASEparameters=124M, task-specific fine-tuning=false2022.11 | — | 64.9 | — | |
| Uni-Perceiver LARGEparameters=354M, task-specific fine-tuning=false2022.11 | — | 67.8 | — | |
| Uni-Perceiver-MoEEvaluation protocol=Zero-Shot, Model size=base-size2023.01 | — | 64.6 | — | |
| Uni-Perceiver-MoEEvaluation protocol=Fine-Tune, Model size=base-size2023.01 | — | 70.5 | — | |
| Uni-Perceiver-MOE BASEparameters=167M, task-specific fine-tuning=false2022.11 | — | 64.6 | — | |
| Uni-Perceiver-MOE LARGEparameters=505M, task-specific fine-tuning=false2022.11 | — | 67.9 | — | |
| Uni-Perceiver-v2 BASEparameters=308M, task-specific fine-tuning=false2022.11 | — | 71.8 | — | |
| Uni-Perceiver-v2 LARGEparameters=446M, task-specific fine-tuning=false2022.11 | — | 75 | — |