Text-to-Video Retrieval on MSVD (test)
2,030R@1JEMC
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| JEMCProtocol=Fine-Tuned2020.03 | 2,030 | 4,780 | 6,110 | — | — | — | 6 | — | — | — | |
| Soft Max MarginProtocol=Fine-Tuned2020.03 | 2,030 | 4,897 | 6,326 | — | — | — | 6 | — | — | — | |
| CEProtocol=Fine-Tuned, Note=use extra labeled data in the form of pre-trained semantic embeddings2020.03 | 1,980 | 4,900 | 6,380 | — | — | — | 6 | — | — | — | |
| HTM-PT*Protocol=Fine-Tuned2020.03 | 1,552 | 4,093 | 5,570 | — | — | — | 8 | — | — | — | |
| Soft Max MarginProtocol=Zero-Shot2020.03 | 1,366 | 3,570 | 4,774 | — | — | — | 12 | — | — | — | |
| HTM-no-PTProtocol=Fine-Tuned2020.03 | 1,300 | 3,743 | 5,241 | — | — | — | 10 | — | — | — | |
| HTM-PT*Protocol=Zero-Shot2020.03 | 1,286 | 3,306 | 4,583 | — | — | — | 13 | — | — | — | |
| CART2024.06 | 63.6 | 87.9 | 92.8 | 73.6 | — | — | — | — | — | — | |
| VidVec-OModel Size=7B, Optimization Data=60K text-only in-context pairs2026.02 | 60.8 | 84.9 | 90.1 | — | — | — | — | — | — | — | |
| InternVideo2 6Bretrieval-stage=dual, zero-shot=true, parameters=6B2026.01 | 59.3 | 84.4 | — | — | — | — | — | — | — | — | |
| InternVideo2_s2-6Bzero-shot=true2026.06 | 59.3 | 84.4 | 89.6 | — | — | — | — | — | — | — | |
| VIRTUE v2 7Bretrieval-stage=dual, zero-shot=true, parameters=7B2026.01 | 57.8 | 82.2 | — | — | — | — | — | — | — | — | |
| MDMMT-22022.03 | 56.8 | 83.1 | 89.2 | — | 1 | 8.8 | — | — | — | — | |
| Side4VideoBackbone=E/14, Pretrain=LAION-2B, Protocol=Frozen backbone2023.11 | 56.1 | 81.7 | 88.8 | — | 1 | 8.4 | — | — | — | — | |
| Side4VideoBackbone=ViT-E/14, Tuning=Frozen backbone, Pretrain=LAION-2B2023.11 | 56.1 | 81.7 | 88.8 | — | 1 | 8.4 | — | — | — | — | |
| LamRAModel Size=7B, Training Scale=Large vision-text scale2026.02 | 55.7 | 81.7 | 88 | — | — | — | — | — | — | — | |
| Side4VideoBackbone=L/14, Pretrain=CLIP-400M, Protocol=Frozen backbone2023.11 | 54.9 | 82.1 | 89.3 | — | 1 | 7.5 | — | — | — | — | |
| Side4VideoBackbone=ViT-L/14, Tuning=Frozen backbone, Pretrain=CLIP-400M2023.11 | 54.9 | 82.1 | 89.3 | — | 1 | 7.5 | — | — | — | — | |
| VAST, HowToCaption-finetunedV. Encoder=ViT-G2023.10 | 54.8 | 80.9 | 87.2 | — | — | — | 1 | — | — | — | |
| BridgeFormerEvaluation Protocol=Fine-tuned2022.01 | 54.4 | 82.8 | 89.4 | — | 1 | 6.1 | — | — | — | — | |
| MILESEvaluation Protocol=Fine-tuning2022.04 | 53.9 | 83.5 | 90.2 | — | 1 | — | — | — | — | — | |
| ELVA-7Bzero-shot=true2026.06 | 53.9 | 80.7 | 87.9 | — | — | — | — | — | — | — | |
| B3Model Size=7B, Training Scale=Large vision-text scale2026.02 | 53.8 | 80 | 86.7 | — | — | — | — | — | — | — | |
| LanguageBindModality=Video2024.06 | 53.5 | 80.5 | 87.5 | 60.6 | — | — | — | — | — | — | |
| LamRA 7Bretrieval-stage=dual, zero-shot=true, parameters=7B2026.01 | 52.4 | 79.8 | — | — | — | — | — | — | — | — | |
| LamRA-7Bzero-shot=true2026.06 | 52.4 | 79.8 | 87 | — | — | — | — | — | — | — | |
| VLM2VecModel Size=7B, Training Scale=Large vision-text scale2026.02 | 52.1 | 78.4 | 85.8 | — | — | — | — | — | — | — | |
| UniME-V2Model Size=7B, Training Scale=Large vision-text scale2026.02 | 52.1 | 77.5 | 84.4 | — | — | — | — | — | — | — | |
| HAT-VTRZero-shot=true2026.02 | 52.09 | 77.16 | 85.82 | — | — | — | — | — | — | — | |
| BridgeFormerevaluation_protocol=fine-tuning2022.01 | 52 | 82.8 | 90 | — | 1 | — | — | — | — | — | |
| BridgeFormerEvaluation Protocol=fine-tuning2022.01 | 52 | 82.8 | 90 | — | 1 | — | — | — | — | — | |
| Cap4 Video2024.06 | 51.8 | 80.8 | 88.3 | — | — | — | — | — | — | — | |
| Cap4VideoPretrain=CLIP-400M, Protocol=Full Fine-tuning2023.11 | 51.8 | 80.8 | 88.3 | — | 1 | 8.3 | — | — | — | — | |
| Cap4VideoTuning=Full Fine-tuning2023.11 | 51.8 | 80.8 | 88.3 | — | 1 | 8.3 | — | — | — | — | |
| STANPretrain=CLIP-400M, Protocol=Full Fine-tuning2023.11 | 51.5 | 80.4 | 88.5 | — | 1 | — | — | — | — | — | |
| STANTuning=Full Fine-tuning2023.11 | 51.5 | 80.4 | 88.5 | — | 1 | — | — | — | — | — | |
| OA-Trans2021.12 | 51.4 | 82.3 | 88 | — | 2 | — | — | — | — | — | |
| CenterCLIPBackbone=ViT-B/16, Clustering Strategy=k-medoids++, Clustering Configuration=B6-4, 160, MeM. (GB)=17.6, Speed (ms)=86.52022.05 | 50.6 | 80.3 | 88.4 | — | 1 | 8.4 | — | — | — | — | |
| VAST†V. Encoder=ViT-G2023.10 | 50.6 | 76.2 | 84.1 | — | — | — | 1 | — | — | — | |
| UNITEModel Size=7B, Training Scale=Large vision-text scale2026.02 | 50.4 | 78.2 | 86.4 | — | — | — | — | — | — | — | |
| CLIP2TVbackbone=ViT-B/162021.11 | 50.2 | 79.8 | 87.9 | — | 1 | 8.6 | — | — | — | — | |
| LanguageBindretrieval-stage=single, zero-shot=true, architecture=CLIP-based2026.01 | 50 | 77.7 | — | — | — | — | — | — | — | — | |
| Unmasked Teacher-17MV. Encoder=ViT-L2023.10 | 49.9 | 77.7 | 85.3 | — | — | — | — | — | — | — | |
| UMT-Lzero-shot=true2026.06 | 49.9 | 77.7 | 85.3 | — | — | — | — | — | — | — | |
| RAPType=Adapter, DSL post-processing=true2024.05 | 49.8 | 86.1 | — | — | — | 9.7 | — | — | — | — | |
| VIRTUE-Embed 7Bretrieval-stage=single, zero-shot=true, architecture=MLLM-based, parameters=7B2026.01 | 49.8 | 77.6 | — | — | — | — | — | — | — | — | |
| CAMoE2022.03 | 49.8 | 79.2 | 87 | — | — | 9.4 | — | — | — | — | |
| DREAMYear=-2026.06 | 49.7 | 79.1 | 87.3 | — | 2 | 8.5 | — | — | — | — | |
| CLIP4clip (meanP)Backbone=ViT-B/16, MeM. (GB)=25.7, Speed (ms)=59.62022.05 | 49.6 | 79.5 | 88 | — | 2 | 8.6 | — | — | — | — | |
| MMRet-v1.5Model Size=7B, Training Scale=Large vision-text scale2026.02 | 49.5 | 76.6 | 83.5 | — | — | — | — | — | — | — | |
| VLM2Vec 7Bretrieval-stage=single, zero-shot=true, architecture=MLLM-based, parameters=7B2026.01 | 49.5 | — | — | — | — | — | — | — | — | — | |
| ViCLIPzero-shot=true2026.06 | 49.1 | — | — | — | — | — | — | — | — | — | |
| Side4VideoBackbone=B/16, Pretrain=CLIP-400M, Protocol=Frozen backbone2023.11 | 49 | 78.5 | 86.7 | — | 2 | 9.1 | — | — | — | — | |
| Side4VideoBackbone=ViT-B/16, Tuning=Frozen backbone, Pretrain=CLIP-400M2023.11 | 49 | 78.5 | 86.7 | — | 2 | 9.1 | — | — | — | — | |
| BLIP*, HowToCaption-finetunedV. Encoder=ViT-B2023.10 | 49 | 76.2 | 84.2 | — | — | — | 2 | — | — | — | |
| BLIP, COCO-finetunedV. Encoder=ViT-B2023.10 | 48.7 | 75.8 | 83.1 | — | — | — | 2 | — | — | — | |
| VideoAlignerYear=20262026.06 | 48.5 | 78.2 | 86.5 | — | — | 8.6 | — | — | — | — | |
| BridgeFormerEvaluation Protocol=Zero-shot2022.01 | 48.4 | 76.4 | 85.8 | — | 2 | 7.4 | — | — | — | — | |
| DRL2024.01 | 48.3 | 79.1 | 87.3 | — | — | — | — | 71.6 | — | — | |
| VLM2Vec-V2Model Size=7B, Training Scale=Large vision-text scale2026.02 | 48.2 | 75.4 | 83.1 | — | — | — | — | — | — | — | |
| QB-Norm+CLIP2Video2022.03 | 48 | 77.9 | 86.2 | — | 2 | — | — | — | — | — | |
| LSDOYear=20252026.06 | 48 | 78.1 | 86.7 | — | 2 | 8.8 | — | — | — | — | |
| NeighborRetr2025.03 | 47.9 | 77.3 | 86 | — | 2 | 9.2 | — | — | — | — | |
| HTVRYear=20252026.06 | 47.7 | 75.6 | 85.5 | — | 2 | 9.9 | — | — | — | — | |
| CenterCLIPBackbone=ViT-B/32, Clustering Strategy=k-medoids++, Clustering Configuration=B6-4, 49, MeM. (GB)=15.0, Speed (ms)=22.92022.05 | 47.6 | 77.5 | 86 | — | 2 | 9.8 | — | — | — | — | |
| QB-NormYear=20222026.06 | 47.6 | 77.6 | 86.1 | — | 2 | — | — | — | — | — | |
| TCRZero-shot=true2026.02 | 47.46 | 72.99 | 83.43 | — | — | — | — | — | — | — | |
| UMTVariant=B2024.06 | 47.4 | 76.8 | 84 | — | — | — | — | — | — | — | |
| CenterCLIPBackbone=ViT-B/32, Clustering Strategy=spectral, Clustering Configuration=B6-4, 49, MeM. (GB)=14.9, Speed (ms)=40.82022.05 | 47.4 | 76.5 | 85.2 | — | 2 | 9.7 | — | — | — | — | |
| UCOFIAMethod Category=CLIP-based Methods2023.09 | 47.4 | 77.6 | — | — | — | 9.6 | — | — | — | — | |
| CM AdapterPretrain=CLIP-400M, Protocol=Frozen backbone2023.11 | 47.4 | 76.6 | 85 | — | 2 | 10.2 | — | — | — | — | |
| CM AdapterTuning=Frozen backbone2023.11 | 47.4 | 76.6 | 85 | — | — | 10.2 | — | — | — | — | |
| CenterCLIPBackbone=ViT-B/32, Clustering Strategy=k-medoids++, Clustering Configuration=B6-3, 49, MeM. (GB)=14.2, Speed (ms)=22.92022.05 | 47.3 | 76.8 | 85.6 | — | 2 | 9.9 | — | — | — | — | |
| CenterCLIPBackbone=ViT-B/32, Clustering Strategy=spectral, Clustering Configuration=B6-3, 49, MeM. (GB)=14.2, Speed (ms)=43.62022.05 | 47.3 | 76.9 | 86 | — | 2 | 9.7 | — | — | — | — | |
| CenterCLIP2023.09 | 47.3 | 76.9 | 86 | — | 2 | 9.7 | — | — | — | — | |
| PAU2023.09 | 47.3 | 77.4 | 85.5 | — | 2 | 9.6 | — | — | — | — | |
| X-Pool2022.03 | 47.2 | 77.4 | 86 | — | 2 | 9.3 | — | — | — | — | |
| X-PoolMethod Category=CLIP-based Methods2023.09 | 47.2 | 77.4 | — | — | — | 9.3 | — | — | — | — | |
| X-Pool2024.01 | 47.2 | 77.4 | 86 | — | — | — | — | 70.2 | — | — | |
| XPool2023.09 | 47.2 | 77.4 | 86 | — | 2 | 9.3 | — | — | — | — | |
| CLIPZero-shot=true2026.02 | 47.16 | 71.64 | 81.79 | — | — | — | — | — | — | — | |
| TentZero-shot=true2026.02 | 47.16 | 71.49 | 82.09 | — | — | — | — | — | — | — | |
| SARZero-shot=true2026.02 | 47.16 | 71.64 | 82.09 | — | — | — | — | — | — | — | |
| X-CLIPMethod Category=CLIP-based Methods2023.09 | 47.1 | 77.8 | — | — | — | 9.5 | — | — | — | — | |
| Align and TellYear=20232026.06 | 47.1 | 77 | 85.6 | — | 2 | — | — | — | — | — | |
| READZero-shot=true2026.02 | 47.01 | 71.79 | 81.34 | — | — | — | — | — | — | — | |
| CLIP2Video2024.06 | 47 | 76.8 | 85.9 | — | — | — | — | — | — | — | |
| CLIP2Video2021.11 | 47 | 76.8 | 85.9 | — | 2 | 9.6 | — | — | — | — | |
| CLIP2TVsequential_distillation=true2021.11 | 47 | 76.5 | 85.1 | — | 2 | 10.1 | — | — | — | — | |
| EERCF2024.01 | 47 | 77.5 | 85.4 | — | — | — | — | 70 | — | — | |
| CLIP2Video2023.09 | 47 | 76.8 | 85.9 | — | 2 | 9.6 | — | — | — | — | |
| CLIP2TV2023.09 | 47 | 76.5 | 85.1 | — | 2 | 10.1 | — | — | — | — | |
| CLIP2TVBackbone=ViT-B/162023.09 | 47 | 76.5 | 85.1 | — | 2 | — | — | — | 10.1 | 208.6 | |
| CLIP2Video2022.03 | 47 | 76.8 | 85.9 | — | 2 | 9.6 | — | — | — | — | |
| CAMOEMethod Category=CLIP-based Methods2023.09 | 46.9 | 76.1 | — | — | — | 9.8 | — | — | — | — | |
| Ours_allBackbone=ViT-B/16, variant=all2023.09 | 46.9 | 76.8 | 85.6 | — | 2 | — | — | — | 9.7 | 209.3 | |
| X-CLIPvariant=dagger symbol marked in paper2024.01 | 46.7 | 76.8 | 85.5 | — | — | — | — | 69.7 | — | — | |
| Diffusion2025.03 | 46.6 | 75.9 | 84.1 | — | 2 | 15.7 | — | — | — | — | |
| EATAZero-shot=true2026.02 | 46.57 | 72.84 | 83.13 | — | — | — | — | — | — | — | |
| DRLBackbone=ViT-B/16, re-trained=true2023.09 | 46.5 | 76.3 | 85 | — | 2 | — | — | — | 10.7 | 207.8 |