Image-to-Text Retrieval on COCO 5K (test)
81.9R@1Uncompressed
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| UncompressedReduce Ratio=None, GFLOPS=153.2, Base Model=BLIP2024.03 | 81.9 | 95.4 | 97.8 | |
| UncompressedModel=BLIP, Pruning Mode=/, Pruning Ratio=/, GFLOPs=91.652026.04 | 81.9 | 95.4 | — | |
| UncompressedModel=BLIP, Pruning Mode=None, Pruning Ratio=None, GFLOPs=91.652026.04 | 81.9 | 95.4 | — | |
| MADTPReduce Ratio=0.5, GFLOPS=87.4, Base Model=BLIP2024.03 | 79.1 | 94.2 | 97.2 | |
| UPopReduce Ratio=0.5, GFLOPS=88.3, Base Model=BLIP2024.03 | 77.4 | 93.4 | 97 | |
| CoMPModel=BLIP, Pruning Mode=C, Pruning Ratio=0.7, GFLOPs=30.082026.04 | 76.2 | 92.4 | — | |
| CoMPModel=BLIP, Pruning Mode=C, Pruning Ratio=0.7, GFLOPs=30.082026.04 | 76.2 | 92.4 | — | |
| MADTPModel=BLIP, Pruning Mode=T, Pruning Ratio=0.7, GFLOPs=30.692026.04 | 73.9 | 91.5 | — | |
| MADTPModel=BLIP, Pruning Mode=T, Pruning Ratio=0.7, GFLOPs=30.692026.04 | 73.9 | 91.5 | — | |
| UPopModel=BLIP, Pruning Mode=P, Pruning Ratio=0.7, GFLOPs=34.332026.04 | 73 | 91.9 | — | |
| UPopModel=BLIP, Pruning Mode=P, Pruning Ratio=0.7, GFLOPs=34.332026.04 | 73 | 91.9 | — | |
| MADTPBackbone=CLIP, Reduce Ratio=0.5, GFLOPS=190.22024.03 | 72.7 | 91.8 | 96.1 | |
| VACSRBackbone=CLIP ViT-B/162026.05 | 72.2 | 91.1 | 95.4 | |
| UncompressedBackbone=CLIP, Reduce Ratio=/, GFLOPS=395.72024.03 | 71.5 | 90.8 | 95.1 | |
| UncompressedModel=CLIP, Pruning Mode=/, Pruning Ratio=/, GFLOPs=197.82026.04 | 71.5 | 90.8 | — | |
| UncompressedModel=CLIP, Pruning Mode=None, Pruning Ratio=None, GFLOPs=197.82026.04 | 71.5 | 90.8 | — | |
| MADTPReduce Ratio=0.75, GFLOPS=50.2, Base Model=BLIP2024.03 | 71.2 | 90 | 94 | |
| UPopBackbone=CLIP, Reduce Ratio=0.5, GFLOPS=196.32024.03 | 70.8 | 90.8 | 95.2 | |
| SLQ (InternVL3-8B)Model category=Ours, Backbone=InternVL3, Scale=8B2026.04 | 69.6 | 89.1 | 93.8 | |
| CoMPModel=CLIP, Pruning Mode=C, Pruning Ratio=0.75, GFLOPs=45.12026.04 | 68.7 | 88.9 | — | |
| CoMPModel=CLIP, Pruning Mode=C, Pruning Ratio=0.75, GFLOPs=45.12026.04 | 68.7 | 88.9 | — | |
| PCME++Backbone=CLIP ViT-B/162026.05 | 68.7 | 90.1 | 95 | |
| VLM2VEC-7BModel category=MLLM-based, Scale=7B2026.04 | 68.5 | 88.4 | 93.4 | |
| VACSRBackbone=CLIP ViT-B/322026.05 | 66.5 | 88.3 | 93.9 | |
| MADTPBackbone=CLIP, Reduce Ratio=0.75, GFLOPS=92.42024.03 | 66.2 | 88.4 | 93.7 | |
| MADTPModel=CLIP, Pruning Mode=T, Pruning Ratio=0.75, GFLOPs=46.22026.04 | 66.2 | 88.4 | — | |
| MADTPModel=CLIP, Pruning Mode=T, Pruning Ratio=0.75, GFLOPs=46.22026.04 | 66.2 | 88.4 | — | |
| PCMEBackbone=CLIP ViT-B/162026.05 | 65.3 | 89.2 | 94.5 | |
| FuseMixTraining Data Size=5M, Regime=low-data regime, Variant=(D,E)2023.12 | 64.3 | 86.2 | 92.1 | |
| SLQ (Qwen3VL-4B)Model category=Ours, Backbone=Qwen3VL, Scale=4B2026.04 | 64.3 | 85.6 | 91.2 | |
| 3TTraining Data Size=5B, Regime=internet-scale2023.12 | 64.1 | — | — | |
| BLIP ViT-LModel category=Dual-Encoder, Backbone=ViT-L2026.04 | 63.5 | 86.5 | 92.5 | |
| UPopReduce Ratio=0.75, GFLOPS=50.2, Base Model=BLIP2024.03 | 62.9 | 86.2 | 92.3 | |
| FuseMixTraining Data Size=5M, Regime=low-data regime, Variant=(D,B)2023.12 | 62.7 | 86.4 | 92.7 | |
| SLQ (Qwen3VL-2B)Model category=Ours, Backbone=Qwen3VL, Scale=2B2026.04 | 62.7 | 84.4 | 90.7 | |
| PCME++Backbone=CLIP ViT-B/322026.05 | 62.1 | 87 | 93 | |
| E5-V-7BModel category=MLLM-based, Scale=7B2026.04 | 62 | 83.6 | 89.7 | |
| FILIPTraining Data Size=300M, Regime=internet-scale2023.12 | 61.3 | 84.3 | 90.4 | |
| SLQ (InternVL3-1B)Model category=Ours, Backbone=InternVL3, Scale=1B2026.04 | 61.2 | 84.9 | 91.8 | |
| FLAMEModel category=Dual-Encoder2026.04 | 60.5 | 82.9 | 89.3 | |
| PCMEBackbone=CLIP ViT-B/322026.05 | 59.9 | 85.8 | 92.3 | |
| DAABackbone=CLIP ViT-B/322026.05 | 59.8 | 85 | 92 | |
| LITTraining Data Size=4B, Regime=internet-scale2023.12 | 59.5 | — | — | |
| FuseMixTraining Data Size=5M, Regime=low-data regime, Variant=(U,B)2023.12 | 59.1 | 83.4 | 90.3 | |
| FuseMixTraining Data Size=5M, Regime=low-data regime, Variant=(U,E)2023.12 | 59.1 | 83.9 | 91 | |
| ALIGNTraining Data Size=1B, Regime=internet-scale2023.12 | 58.6 | 83 | 89.7 | |
| CLIPTraining Data Size=400M, Regime=internet-scale2023.12 | 58.4 | 81.5 | 88.1 | |
| CLIP ViT-LModel category=Dual-Encoder, Backbone=ViT-L2026.04 | 58.1 | 81 | 87.8 | |
| P2RMBackbone=CLIP ViT-B/162026.05 | 56.8 | 84.3 | 91.5 | |
| P2RMBackbone=CLIP ViT-B/322026.05 | 56.6 | 83.5 | 90.9 | |
| ReasonCLIP (Stage 1)Model Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=+42M, Base Model=CLIP2026.06 | 56.2 | 78.8 | — | |
| UPopBackbone=CLIP, Reduce Ratio=0.75, GFLOPS=105.92024.03 | 56.1 | 82.4 | 90.2 | |
| UPopModel=CLIP, Pruning Mode=P, Pruning Ratio=0.75, GFLOPs=57.92026.04 | 56.1 | 82.4 | — | |
| UPopModel=CLIP, Pruning Mode=P, Pruning Ratio=0.75, GFLOPs=57.92026.04 | 56.1 | 82.4 | — | |
| DataCompModel Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=1.4B2026.06 | 53.4 | 77.5 | — | |
| OpenCLIPModel Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=0.4B2026.06 | 52.3 | 76.3 | — | |
| ReasonCLIP (Stage 2)Model Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=+16M, Base Model=CLIP2026.06 | 52.3 | 77.2 | — | |
| MetaCLIPModel Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=0.4B2026.06 | 51.8 | 76.4 | — | |
| CLIP ViT-BModel category=Dual-Encoder, Backbone=ViT-B2026.04 | 51 | 74.9 | 83.5 | |
| CLIPModel Scale=ViT Base (86M), Resolution=224/32, Training Data Scale=0.4B2026.06 | 50 | 75 | — | |
| FuseMixTraining Data Size=3M, Regime=low-data regime, Variant=(D,B)2023.12 | 42.3 | 68.4 | 78.9 | |
| CLIPTraining Data Size=3M, Regime=low-data regime2023.12 | 36.2 | 64.3 | 80.1 | |
| DAABackbone=CLIP ViT-B/162026.05 | 24.3 | 49.9 | 62.7 |