Image-to-Text Retrieval on Flickr30k
99R@1Qwen-2.5
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen-2.52025.12 | 99 | 100 | 100 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Finetuning2026.02 | 98.9 | 100 | 100 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Finetuning2026.02 | 98.8 | 100 | 100 | — | |
| BaselineBackbone=SigLip 2 (ViT-B/16), CR=0, Setting=Finetuning2026.02 | 98.7 | 100 | 100 | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Finetuning2026.02 | 98.5 | 100 | 100 | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Finetuning2026.02 | 98.1 | 100 | 100 | — | |
| BLIP-22025.12 | 98 | 100 | 100 | — | |
| MDSE2025.12 | 98 | 100 | 100 | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Finetuning2026.02 | 97.8 | 99.9 | 100 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=87.5%, Setting=Finetuning2026.02 | 97.7 | 99.9 | 100 | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Finetuning2026.02 | 97.5 | 99.8 | 100 | — | |
| BLIPTraining Protocol=full fine-tuning, # Tunable=223M2023.05 | 97.3 | 99.9 | 100 | — | |
| Florence#param=893M, #data=900M, tuning_mode=fine-tuning (FT100%), resolution=3842022.06 | 97.2 | — | — | — | |
| AuroraTraining Protocol=frozen backbone, # Tunable=0.2M, Rank (r)=1282023.05 | 97.2 | 100 | 100 | — | |
| UniAdapterTraining Protocol=frozen backbone, # Tunable=4.6M, Rank (r)=1282023.05 | 97.1 | 100 | 100 | — | |
| UniAdapterTraining Protocol=frozen backbone, # Tunable=18.8M, Rank (r)=5122023.05 | 97.1 | 99.9 | 100 | — | |
| BLIP-2Evaluation Protocol=Zero-shot2024.10 | 96.9 | 100 | 100 | — | |
| AuroraTraining Protocol=frozen backbone, # Tunable=0.1M, Rank (r)=642023.05 | 96.8 | 100 | 100 | — | |
| FG-CLIP 2Backbone=ViT-L/16, Evaluation Protocol=Zero-shot2025.10 | 96.6 | — | — | — | |
| LoRATraining Protocol=frozen backbone, # Tunable=10.6M, Rank (r)=322023.05 | 96.2 | 99.7 | 99.8 | — | |
| ALBEFTraining Protocol=full fine-tuning, # Tunable=210M2023.05 | 95.9 | 99.8 | 100 | — | |
| FG-CLIP2Backbone=ViT-So/162026.02 | 95.9 | — | — | — | |
| FG-CLIP 2Backbone=ViT-So/16, Evaluation Protocol=Zero-shot2025.10 | 95.9 | — | — | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=87.5%, Setting=Finetuning2026.02 | 95.6 | 99.6 | 100 | — | |
| ALIGN#param=820M, #data=1.8B, tuning_mode=fine-tuning (FT100%), resolution=2892022.06 | 95.3 | — | — | — | |
| ALIGNTraining Protocol=full fine-tuning, # Tunable=820M2023.05 | 95.3 | 99.8 | 100 | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=87.5%, Setting=Finetuning2026.02 | 95.3 | 99.4 | 100 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=93.75%, Setting=Finetuning2026.02 | 95.3 | 99.5 | 100 | — | |
| MADTPGFLOPs=74.52024.03 | 95.1 | 99.5 | 99.7 | — | |
| BEIT3Model scale=0.7B, Zero-shot evaluation=true2024.01 | 94.9 | 99.9 | 100 | — | |
| Uni-Perceiver-L#param=354M, #data=44.1M, tuning_mode=fine-tuning (FT100%)2022.06 | 94.7 | — | — | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Zero-Shot2026.02 | 94.6 | 99.5 | 100 | — | |
| BaselineBackbone=SigLip 2 (ViT-B/16), CR=0, Setting=Zero-Shot2026.02 | 94.4 | 99.5 | 100 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Zero-Shot2026.02 | 94.4 | 99.5 | 100 | — | |
| AMoEHead=Ensemble, Resolution=512x512, Parameters=0.3B-0.6B2025.12 | 94.3 | — | — | — | |
| SigLIP 2Backbone=ViT-L/16, Evaluation Protocol=Zero-shot2025.10 | 94.3 | — | — | — | |
| AAPEEvaluation Protocol=Train or Finetune, Note=Learned on COCO2024.10 | 94.2 | 99.3 | 99.7 | — | |
| ObjEmbedParameters=2B2026.02 | 94.2 | — | — | — | |
| ObjEmbedParameters=4B2026.02 | 94.2 | — | — | — | |
| Uni-Perceiver-L + Conditional MoEs#param=505M, #data=44.1M, tuning_mode=fine-tuning (FT100%)2022.06 | 94.1 | — | — | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Zero-Shot2026.02 | 94.1 | 99.3 | 100 | — | |
| SigLIP2Backbone=ViT-So/162026.02 | 94.1 | — | — | — | |
| FG-CLIP 2Backbone=ViT-B/16, Evaluation Protocol=Zero-shot2025.10 | 94.1 | — | — | — | |
| SigLIP 2Backbone=ViT-So/16, Evaluation Protocol=Zero-shot2025.10 | 94.1 | — | — | — | |
| UPopGFLOPs=91.02024.03 | 94 | 99.5 | 99.7 | — | |
| EVAModel scale=5.0B, Zero-shot evaluation=true2024.01 | 93.9 | 99.4 | 99.8 | — | |
| MADTPBackbone=CLIP, GFLOPs=178.82024.03 | 93.9 | 99.5 | 99.8 | — | |
| FG-CLIPBackbone=ViT-L/14, Evaluation Protocol=Zero-shot2025.10 | 93.7 | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, tuning_mode=fine-tuning (FT100%)2022.06 | 93.6 | — | — | — | |
| RADIOv2.5-HHead=Ensemble, Resolution=512x512, Parameters=0.6B2025.12 | 93.5 | — | — | — | |
| UMG-CLIPModel scale=0.4B, Zero-shot evaluation=true2024.01 | 93.4 | 99.5 | 99.9 | — | |
| UMG-CLIPBackbone=ViT-L/14, Evaluation Protocol=Zero-shot2025.10 | 93.4 | — | — | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=93.75%, Setting=Finetuning2026.02 | 93.3 | 99 | 99.5 | — | |
| UPopBackbone=CLIP, GFLOPs=201.12024.03 | 93.2 | 99.4 | 99.8 | — | |
| LlipEvaluation Protocol=Zero-shot2024.10 | 93.2 | 99 | 99.4 | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=50%, Setting=Zero-Shot2026.02 | 93.2 | 99 | 100 | — | |
| UMG-CLIPModel scale=5B, Zero-shot evaluation=true2024.01 | 93.1 | 99.7 | 99.9 | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=93.75%, Setting=Finetuning2026.02 | 93.1 | 99.1 | 99.4 | — | |
| RADIOv2.5-LHead=Ensemble, Resolution=512x512, Parameters=0.3B2025.12 | 93.1 | — | — | — | |
| OpenCLIPModel scale=2.5B, Zero-shot evaluation=true2024.01 | 92.9 | 99.3 | 99.8 | — | |
| Qwen3-VL-EmbeddingParameters=2B2026.02 | 92.9 | — | — | — | |
| RADIOv2.5-HHead=CLIP, Resolution=512x512, Parameters=0.6B2025.12 | 92.9 | — | — | — | |
| Uni-Perceiver-B#param=124M, #data=44.1M, tuning_mode=fine-tuning (FT100%)2022.06 | 92.7 | — | — | — | |
| TeachersHead=SigLIP2, Resolution=512x5122025.12 | 92.6 | — | — | — | |
| SigLIP 2Backbone=ViT-B/16, Evaluation Protocol=Zero-shot2025.10 | 92.6 | — | — | — | |
| CoCa-L#param=787M, #data=4.8B, tuning_mode=w/o tuning (WT), resolution=5762022.06 | 92.5 | — | — | — | |
| CoCaModel scale=2.1B, Zero-shot evaluation=true2024.01 | 92.5 | 99.5 | 99.9 | — | |
| Arbitrary Ratio Feature Compression (ARFC)Backbone=SigLip 2 (ViT-B/16), CR=87.5%, Setting=Zero-Shot2026.02 | 92.5 | 98.9 | 99.9 | — | |
| RADIOv2.5-LHead=CLIP, Resolution=512x512, Parameters=0.3B2025.12 | 92.5 | — | — | — | |
| Uni-Perceiver-L + Conditional MoEs#param=505M, #data=44.1M, tuning_mode=prompt tuning (PT1%)2022.06 | 92.4 | — | — | — | |
| Ensemble TeacherTraining Dataset=DataComp-1B [18], OpenAI-400M [47], Image Encoder=ViT-L/142023.11 | 92.3 | — | — | — | |
| MobileCLIP-B (LT)Training Dataset=DataCompDR-1B, Image Encoder=ViT-B/162023.11 | 92.3 | — | — | — | |
| ELIPGFLOPs=93.42024.03 | 92.2 | 99.1 | 99.7 | — | |
| RADIOv2.5-LHead=SigLIP, Resolution=512x512, Parameters=0.3B2025.12 | 92.2 | — | — | — | |
| RADIOv2.5-HHead=SigLIP, Resolution=512x512, Parameters=0.6B2025.12 | 92.2 | — | — | — | |
| Uni-Perceiver-L#param=354M, #data=44.1M, tuning_mode=prompt tuning (PT1%)2022.06 | 92.1 | — | — | — | |
| CrossGETBackbone=CLIP2024.03 | 92.1 | 99.7 | 99.8 | — | |
| MetaCLIP2Backbone=ViT-H/142026.02 | 91.9 | — | — | — | |
| AMoEHead=SigLIP2, Resolution=512x512, Parameters=0.3B-0.6B2025.12 | 91.9 | — | — | — | |
| Meta CLIP 2Backbone=ViT-H/14, Evaluation Protocol=Zero-shot2025.10 | 91.9 | — | — | — | |
| Q-FormerBackbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Zero-Shot2026.02 | 91.6 | 98.7 | 99.7 | — | |
| UME-R1Parameters=2B2026.02 | 91.6 | — | — | — | |
| ToMeGFLOPs=69.82024.03 | 91.5 | 98.8 | 99.4 | — | |
| SigLIPEvaluation Protocol=Zero-shot2024.10 | 91.5 | 98.1 | 99.4 | — | |
| OpenCLIPModel scale=1.3B, Zero-shot evaluation=true2024.01 | 91.4 | 99.2 | 99.6 | — | |
| CoCaModel scale=0.8B, Zero-shot evaluation=true2024.01 | 91.4 | 99.2 | 99.9 | — | |
| MobileCLIP-BEvaluation Protocol=zero-shot2023.11 | 91.4 | 99.1 | 99.9 | — | |
| MobileCLIP-BTraining Dataset=DataCompDR-1B, Image Encoder=ViT-B/162023.11 | 91.4 | — | — | — | |
| UMG-CLIPBackbone=ViT-B/16, Evaluation Protocol=Zero-shot2025.10 | 91.4 | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, tuning_mode=prompt tuning (PT1%)2022.06 | 91.3 | — | — | — | |
| ERNIE-ViL 2.0Zero-shot evaluation=true2024.01 | 91.2 | 99.1 | 99.8 | — | |
| M2-EncoderModel scale=10B, Zero-shot evaluation=true2024.01 | 91.2 | 99.2 | 99.6 | — | |
| ViCLIP-B/16Evaluation Protocol=zero-shot2023.11 | 91.1 | 98.5 | 99.7 | — | |
| VeCLIP-B/16Training Dataset=WIT-200M, Image Encoder=ViT-B/162023.11 | 91.1 | — | — | — | |
| AutoEncoder-VAEBackbone=SigLip 2 (ViT-B/16), CR=75%, Setting=Zero-Shot2026.02 | 91.1 | 98.5 | 99.7 | — | |
| GMEParameters=7B2026.02 | 91.1 | — | — | — | |
| Uni-Perceiver-B#param=124M, #data=44.1M, tuning_mode=prompt tuning (PT1%)2022.06 | 91 | — | — | — | |
| AMoEHead=DINOv3, Resolution=512x512, Parameters=0.3B-0.6B2025.12 | 91 | — | — | — | |
| Florence2025.12 | 91 | 99 | — | — | |
| Florence#param=893M, #data=900M, tuning_mode=w/o tuning (WT), resolution=3842022.06 | 90.9 | — | — | — |