Text-to-Video Retrieval on MSVD
77.08R@1Gemini Embedding 2
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Gemini Embedding 2Setting=ZS2026.05 | 77.08 | — | — | — | — | — | — | |
| UMTtwo-stage candidate re-ranking=true2024.03 | 71.9 | 94.5 | 97.8 | — | — | — | — | |
| mPLUG2+Quantization=false, Weight precision=32-bit floating point, Text decoder=LLaMa2025.12 | 69.1 | 84.5 | 93.8 | — | — | — | — | |
| OmniRetriever-7BSetting=ZS2026.05 | 66.88 | — | — | — | — | — | — | |
| mPLUG2+Quantization=true, Weight precision=8-bit integer, Text decoder=LLaMa2025.12 | 66.8 | 82.7 | 90.3 | — | — | — | — | |
| mPLUG2Quantization=false, Weight precision=32-bit floating point, Text decoder=BERT2025.12 | 65.3 | 82 | 90.6 | — | — | — | — | |
| mPLUG2Quantization=true, Weight precision=8-bit integer, Text decoder=BERT2025.12 | 62.1 | 77.3 | 87.9 | — | — | — | — | |
| InternVideo2s2-6BFinetuning=true2024.03 | 61.4 | — | — | — | — | — | — | |
| VidVecV-T Pairs=Text-Only2026.02 | 60.9 | — | — | — | — | — | — | |
| PEAV-BV-Enc Params=0.4B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 60.8 | — | — | — | — | — | — | |
| PEAV-LV-Enc Params=0.5B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 60.8 | — | — | — | — | — | — | |
| PEAV-BV-Enc Params=0.4B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 60.7 | — | — | — | — | — | — | |
| PEAV-LV-Enc Params=0.5B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 60.5 | — | — | — | — | — | — | |
| PEAV-SV-Enc Params=0.3B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 60.1 | — | — | — | — | — | — | |
| PEAV-SV-Enc Params=0.3B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 59.8 | — | — | — | — | — | — | |
| PE-Core-GV-T Pairs=22M2026.02 | 59.7 | — | — | — | — | — | — | |
| PE-GV-Enc Params=1.9B, Video Data=22M, Resolution=448, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 59.7 | — | — | — | — | — | — | |
| PE-coreGSetting=ZS2026.05 | 59.7 | — | — | — | — | — | — | |
| InternVideo2_S2-6BSetting=Zero-shot2024.12 | 59.3 | 84.4 | 89.6 | — | — | — | — | |
| InternVideo2-6BV-T Pairs=100M2026.02 | 59.3 | — | — | — | — | — | — | |
| InternVideo2_s2-6BEvaluation Protocol=Zero-shot (finetuned with text-video data)2025.07 | 59.3 | 84.4 | 89.6 | — | — | — | — | |
| InternVideo2-6BFinetuning data modality=text-video data2025.03 | 59.3 | 84.4 | 89.6 | — | — | — | — | |
| CLIP2videoParameters=132M, Data=400M, Protocol=Weight Transfer2022.06 | 58.7 | — | — | — | — | — | — | |
| InternVideoTraining Data (M)=210, Matching Paradigm=vision-text contrastive2023.05 | 58.4 | 84.5 | 90.4 | — | — | 233.3 | — | |
| InternVideo#Pairs=646M2023.03 | 58.4 | 84.5 | 90.4 | — | — | — | — | |
| InternVideo(ViT-L)Backbone=ViT-L, dual-softmax inference=true2024.03 | 58.4 | 84.5 | 90.4 | — | — | — | — | |
| InternVideo#Pairs=646M2024.03 | 58.4 | — | — | — | — | — | — | |
| UMT-L#Pairs=25M2023.03 | 58.2 | 83.9 | 89.6 | — | — | — | — | |
| UMT-L#Pairs=25M2024.03 | 58.2 | — | — | — | — | — | — | |
| UMT-LFinetuning=true2024.03 | 58.2 | — | — | — | — | — | — | |
| InternVideo2V-Enc Params=1.0B, Video Data=102M, Resolution=224, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 58.1 | — | — | — | — | — | — | |
| InternVideo2-1BSetting=ZS2026.05 | 58.1 | — | — | — | — | — | — | |
| UMT-L + vid-TLDR#Pairs=25M2024.03 | 57.9 | — | — | — | — | — | — | |
| VLAB_GTraining Data (M)=262023.05 | 57.5 | 83.6 | 89.9 | — | — | 231.1 | — | |
| HAT-VTRBase Model=Xpool, Adaptation Method=HAT-VTR2026.02 | 57.46 | 85.97 | 92.84 | — | — | — | — | |
| UMT-L#Pairs=17M2023.03 | 57.4 | 83 | 88.5 | — | — | — | — | |
| PE-LV-Enc Params=0.3B, Video Data=22M, Resolution=336, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 57.2 | — | — | — | — | — | — | |
| PE-coreLSetting=ZS2026.05 | 57.2 | — | — | — | — | — | — | |
| XpoolBase Model=Xpool, Adaptation Method=None2026.02 | 55.97 | 85.37 | 91.79 | — | — | — | — | |
| TentBase Model=Xpool, Adaptation Method=Tent2026.02 | 55.97 | 85.52 | 91.79 | — | — | — | — | |
| READBase Model=Xpool, Adaptation Method=READ2026.02 | 55.97 | 85.22 | 91.79 | — | — | — | — | |
| SARBase Model=Xpool, Adaptation Method=SAR2026.02 | 55.97 | 85.52 | 91.79 | — | — | — | — | |
| EATABase Model=Xpool, Adaptation Method=EATA2026.02 | 55.97 | 85.22 | 91.79 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/16, dual-softmax inference=true2024.03 | 55.9 | 82.3 | 88.3 | — | — | — | — | |
| SigLIP2-g-optV-Enc Params=1.1B, Resolution=384, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 55.8 | — | — | — | — | — | — | |
| SigLIP-2-g-optSetting=ZS2026.05 | 55.8 | — | — | — | — | — | — | |
| LanguageBindV-Enc Params=0.3B, Video Data=10M, Resolution=224, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 55.6 | — | — | — | — | — | — | |
| VLAB_LTraining Data (M)=262023.05 | 55.4 | 82.5 | 89.2 | — | — | 227.1 | — | |
| TCRBase Model=Xpool, Adaptation Method=TCR2026.02 | 55.37 | 85.67 | 91.34 | — | — | — | — | |
| HAT-VTRBase Model=CLIP4Clip, Adaptation Method=HAT-VTR2026.02 | 54.78 | 84.33 | 90.45 | — | — | — | — | |
| CLIP4ClipBase Model=CLIP4Clip, Adaptation Method=None2026.02 | 54.63 | 80.9 | 90.9 | — | — | — | — | |
| READBase Model=CLIP4Clip, Adaptation Method=READ2026.02 | 54.63 | 80.9 | 90.9 | — | — | — | — | |
| EATABase Model=CLIP4Clip, Adaptation Method=EATA2026.02 | 54.63 | 81.19 | 90.45 | — | — | — | — | |
| U-MARVELEvaluation Protocol=Zero-shot (finetuned only with text-image data)2025.07 | 54.6 | 80.9 | 87.7 | — | — | — | — | |
| TentBase Model=CLIP4Clip, Adaptation Method=Tent2026.02 | 54.48 | 80.9 | 90.9 | — | — | — | — | |
| SARBase Model=CLIP4Clip, Adaptation Method=SAR2026.02 | 54.48 | 80.9 | 90.9 | — | — | — | — | |
| FluxViT-BToken count=2048, Dual Softmax Loss=true, Zero-shot=true2025.03 | 54.2 | — | — | — | — | — | — | |
| TCRBase Model=CLIP4Clip, Adaptation Method=TCR2026.02 | 54.18 | 80.9 | 90.45 | — | — | — | — | |
| LanguageBind-LZero-shot=true2025.03 | 54.1 | — | — | — | — | — | — | |
| LanguageBind-HZero-shot=true2025.03 | 53.9 | — | — | — | — | — | — | |
| LanguageBindV-T Pairs=10M2026.02 | 53.9 | — | — | — | — | — | — | |
| LanguageBind (CLIP-H/14)Setting=ZS2026.05 | 53.9 | — | — | — | — | — | — | |
| FluxViT-BToken count=2048, Dual Softmax Loss=false, Zero-shot=true2025.03 | 53.8 | — | — | — | — | — | — | |
| UMT-L#Pairs=5M2023.03 | 53.7 | 80.5 | 86.8 | — | — | — | — | |
| SigLIP2-L/16V-Enc Params=0.3B, Resolution=384, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 53.7 | — | — | — | — | — | — | |
| SigLIP-2-L/16Setting=ZS2026.05 | 53.7 | — | — | — | — | — | — | |
| FluxViT-BToken count=1024, Dual Softmax Loss=true, Zero-shot=true2025.03 | 53.4 | — | — | — | — | — | — | |
| NarVid2025.03 | 53.1 | 81.4 | 88.8 | — | — | — | — | |
| U-MARVEL+Evaluation Protocol=Zero-shot (finetuned only with text-image data)2025.07 | 53.1 | 79.9 | 87.1 | — | — | — | — | |
| LLaVE-7BEvaluation Protocol=Zero-shot (finetuned only with text-image data)2025.07 | 52.9 | 80.1 | 87 | — | — | — | — | |
| LLaVE-7BFinetuning data modality=text-image data2025.03 | 52.9 | 80.1 | 87 | — | — | — | — | |
| FluxViT-BToken count=1024, Dual Softmax Loss=false, Zero-shot=true2025.03 | 52.8 | — | — | — | — | — | — | |
| HunYuan_tvrParameters=364M, Data=400M, Protocol=Weight Transfer2022.06 | 52.7 | — | — | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/32, dual-softmax inference=true2024.03 | 52.7 | 80.4 | 87 | — | — | — | — | |
| MoVA2026.07 | 52.6 | 83 | 90.6 | 1 | 7.8 | — | — | |
| OmniVLTraining Data (M)=18, Matching Paradigm=vision-text matching2023.05 | 52.4 | 79.5 | 85.4 | — | — | 217.3 | — | |
| LamRASetting=Zero-shot2024.12 | 52.4 | 79.8 | 87 | — | — | — | — | |
| LamRAEvaluation Protocol=Zero-shot (finetuned only with text-image data)2025.07 | 52.4 | 79.8 | 87 | — | — | — | — | |
| LamRAFinetuning data modality=text-image data2025.03 | 52.4 | 79.8 | 87 | — | — | — | — | |
| Uni-Perceiver-L + Conditional MoEsParameters=505M, Data=44.1M, Protocol=Fine-tuning (100%)2022.06 | 52.3 | — | — | — | — | — | — | |
| FluxViT-BToken count=512, Dual Softmax Loss=true, Zero-shot=true2025.03 | 52.1 | — | — | — | — | — | — | |
| BridgeFormerBackbone=CLIP-ViT-B/162024.03 | 52 | 82.8 | 90 | — | — | — | — | |
| Cap4VideoBackbone=CLIP-ViT-B/162024.03 | 51.8 | 80.8 | 88.3 | — | — | — | — | |
| Cap4VideoVenue=CVPR'232025.03 | 51.8 | 80.8 | 88.3 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/162024.03 | 51.5 | 79.9 | 86.9 | — | — | — | — | |
| EagleNetBackbone=ViT-B/162026.03 | 50.9 | 80.7 | 88.3 | — | — | — | — | |
| Uni-Perceiver-LParameters=354M, Data=44.1M, Protocol=Fine-tuning (100%)2022.06 | 50.8 | — | — | — | — | — | — | |
| UMT-B#Pairs=25M2023.03 | 50.8 | 79.7 | 86.2 | — | — | — | — | |
| UMT-B#Pairs=25M2024.03 | 50.8 | — | — | — | — | — | — | |
| FluxViT-BToken count=512, Dual Softmax Loss=false, Zero-shot=true2025.03 | 50.7 | — | — | — | — | — | — | |
| UMT-B + vid-TLDR#Pairs=25M2024.03 | 50.5 | — | — | — | — | — | — | |
| X-CLIPBackbone=ViT-B/162022.07 | 50.4 | 80.6 | — | — | 8.4 | — | — | |
| X-CLIPMatching Paradigm=vision-text contrastive2023.05 | 50.4 | 80.6 | — | — | — | — | — | |
| X-CLIPSetting=FT2026.05 | 50.4 | — | — | — | — | — | — | |
| PE-coreBSetting=ZS2026.05 | 50.4 | — | — | — | — | — | — | |
| X-CLIP2026.07 | 50.4 | 80.6 | 89.8 | 1 | 8.4 | — | — | |
| Uni-Perceiver-L + Conditional MoEsParameters=505M, Data=44.1M, Protocol=Prompt Tuning (1%)2022.06 | 50.3 | — | — | — | — | — | — | |
| LAVENDERTraining Data (M)=30, Matching Paradigm=vision-text matching2023.05 | 50.1 | 79.6 | 87.2 | — | — | 216.9 | — | |
| LAVENDER#Pairs=30M2023.03 | 50.1 | 79.6 | 87.2 | — | — | — | — | |
| LAVENDER#Pairs=30M2024.03 | 50.1 | — | — | — | — | — | — |