Text-to-Video Retrieval on DiDeMo
32.4R@1TVTS
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| TVTS2022.09 | 32.4 | 59.8 | — | 3 | — | 71.7 | — | — | — | — | |
| Frozen2022.09 | 31 | 59.8 | — | 3 | — | 72.4 | — | — | — | — | |
| ClipBert2022.09 | 20.4 | 48 | — | 6 | — | 60.8 | — | — | — | — | |
| CE2022.09 | 16.1 | 41.1 | — | 8.3 | — | 82.7 | — | — | — | — | |
| HERO2022.09 | 2.1 | — | — | — | — | 11.4 | — | — | — | — | |
| InternVideo2s2-6BFinetuning=true2024.03 | 0.742 | — | — | — | — | — | — | — | — | — | |
| UMT-L#Pairs=25M2024.03 | 0.725 | — | — | — | — | — | — | — | — | — | |
| UMT-L + vid-TLDR#Pairs=25M2024.03 | 0.723 | — | — | — | — | — | — | — | — | — | |
| VASTSample=442M, Modality=Audio/Subtitle2023.05 | 0.72 | 0.89 | — | — | — | 0.914 | — | — | — | — | |
| VASTmodality track=Multi-modal2023.05 | 0.72 | — | — | — | — | — | — | — | — | — | |
| VAST#Pairs=152M2026.03 | 0.72 | — | — | — | — | — | — | — | — | — | |
| UMT-LSample=425M, Modality=Vision-only2023.05 | 0.704 | 0.901 | — | — | — | 0.935 | — | — | — | — | |
| Method [54]modality track=Vision-only2023.05 | 0.704 | — | — | — | — | — | — | — | — | — | |
| UMT-L#Pairs=25M2023.03 | 0.704 | 0.901 | — | — | — | 0.935 | — | — | — | — | |
| UMT-LFinetuning=true2024.03 | 0.704 | — | — | — | — | — | — | — | — | — | |
| UMT-LModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 0.704 | — | — | — | — | 0.935 | — | — | — | — | |
| vid-TLDRModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 0.704 | — | — | — | — | 0.94 | — | — | — | — | |
| UMT-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 0.704 | — | — | — | — | — | — | — | — | — | |
| vid-TLDRModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 0.704 | — | — | — | — | — | — | — | — | — | |
| UMT-Lsetting=finetuning, evaluation_frames=122025.07 | 0.704 | — | — | — | — | — | — | — | — | — | |
| vid-TLDRsetting=finetuning, evaluation_frames=122025.07 | 0.704 | — | — | — | — | — | — | — | — | — | |
| UMT-L#Pairs=25M2026.03 | 0.704 | 0.901 | — | — | — | 0.935 | — | — | — | — | |
| PMRLsetting=finetuning2025.07 | 0.702 | — | — | — | — | — | — | — | — | — | |
| GRAMsetting=finetuning2025.07 | 0.687 | — | — | — | — | — | — | — | — | — | |
| VASTsetting=finetuning2025.07 | 0.684 | — | — | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Protocol=Finetuning2024.12 | 0.673 | — | — | — | — | 0.901 | — | — | — | — | |
| GRAM ModelModality=T-VA, Finetuning=true2024.12 | 0.673 | — | — | — | — | — | — | — | — | — | |
| UMT-L#Pairs=17M2023.03 | 0.666 | 0.899 | — | — | — | 0.937 | — | — | — | — | |
| GRAM ModelModality=T-V, Protocol=Finetuning2024.12 | 0.664 | — | — | — | — | 0.899 | — | — | — | — | |
| GRAM ModelModality=T-V, Finetuning=true2024.12 | 0.664 | — | — | — | — | — | — | — | — | — | |
| VASTModality=T-VA, Protocol=Finetuning2024.12 | 0.656 | — | — | — | — | 0.881 | — | — | — | — | |
| VASTModality=T-VA, Finetuning=true2024.12 | 0.656 | — | — | — | — | — | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/16, dual-softmax inference=true2024.03 | 0.648 | 0.86 | — | — | — | 0.918 | — | — | — | — | |
| UMT-B + vid-TLDR#Pairs=25M2024.03 | 0.641 | — | — | — | — | — | — | — | — | — | |
| UMT-B#Pairs=25M2024.03 | 0.637 | — | — | — | — | — | — | — | — | — | |
| MAMABackbone=CLIP-ViP2026.01 | 0.627 | 0.899 | — | — | — | 0.96 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/32, dual-softmax inference=true2024.03 | 0.622 | 0.846 | — | — | — | 0.9 | — | — | — | — | |
| VidVecV-T Pairs=Text-Only2026.02 | 0.618 | — | — | — | — | — | — | — | — | — | |
| UMT-B#Pairs=25M2023.03 | 0.616 | 0.868 | — | — | — | 0.915 | — | — | — | — | |
| UMTtwo-stage candidate re-ranking=true2024.03 | 0.616 | 0.868 | — | — | — | 0.915 | — | — | — | — | |
| VINDLU#Pairs=25M2023.03 | 0.612 | 0.858 | — | — | — | 0.91 | — | — | — | — | |
| VINDLUtwo-stage candidate re-ranking=true2024.03 | 0.612 | 0.858 | — | — | — | 0.91 | — | — | — | — | |
| VINDLU#Pairs=25M2024.03 | 0.612 | — | — | — | — | — | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/162024.03 | 0.611 | 0.837 | — | — | — | 0.891 | — | — | — | — | |
| UMT-B#Pairs=17M2023.03 | 0.608 | 0.851 | — | — | — | 0.91 | — | — | — | — | |
| VINDLUTraining Data (M)=25, Matching Paradigm=vision-text matching2023.05 | 0.598 | 0.866 | — | — | — | 0.915 | 237.9 | — | — | — | |
| VINDLU-LSample=25M, Modality=Vision-only2023.05 | 0.598 | 0.866 | — | — | — | 0.915 | — | — | — | — | |
| VindLU2026.01 | 0.598 | 0.866 | — | — | — | 0.915 | — | — | — | — | |
| UMT-L#Pairs=5M2023.03 | 0.597 | 0.849 | — | — | — | 0.908 | — | — | — | — | |
| ClusterSTM#Pairs=5M2026.03 | 0.585 | 0.85 | — | — | — | 0.902 | — | — | — | — | |
| UniAlignModality=T-VA, Zero-shot=true, Tuple-level losses=true2026.02 | 0.582 | — | — | — | — | — | — | — | — | — | |
| InternVideoBackbone=ViT-L/142022.12 | 0.579 | — | — | — | — | — | — | — | — | — | |
| InternVideoTraining Data (M)=210, Matching Paradigm=vision-text contrastive2023.05 | 0.579 | 0.824 | — | — | — | 0.889 | 229.2 | — | — | — | |
| InternVideo#Pairs=646M2023.03 | 0.579 | 0.824 | — | — | — | 0.889 | — | — | — | — | |
| InternVideo(ViT-L)Backbone=ViT-L, dual-softmax inference=true2024.03 | 0.579 | 0.824 | — | — | — | 0.889 | — | — | — | — | |
| InternVideo#Pairs=646M2024.03 | 0.579 | — | — | — | — | — | — | — | — | — | |
| InternVideo-LModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 0.579 | — | — | — | — | — | — | — | — | — | |
| InternVideo-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 0.579 | — | — | — | — | — | — | — | — | — | |
| InternVideo2-6BV-T Pairs=100M2026.02 | 0.579 | — | — | — | — | — | — | — | — | — | |
| InternVideo-Lsetting=finetuning, evaluation_frames=122025.07 | 0.579 | — | — | — | — | — | — | — | — | — | |
| VALOR-LSample=433.5M, Modality=Audio/Subtitle2023.05 | 0.576 | 0.833 | — | — | — | 0.888 | — | — | — | — | |
| Method [7]modality track=Multi-modal2023.05 | 0.576 | — | — | — | — | — | — | — | — | — | |
| RTQVideo-language pretraining (PT)=false2023.12 | 0.576 | 0.841 | — | 1 | — | 0.898 | — | — | — | — | |
| Vid2SeqBackbone=CLIP-ViP2026.01 | 0.576 | 0.799 | — | — | — | 0.884 | — | — | — | — | |
| VALOR-Lsetting=finetuning2025.07 | 0.576 | — | — | — | — | — | — | — | — | — | |
| MoVA2026.07 | 0.575 | 0.832 | — | 1 | 4.8 | 0.916 | — | — | — | — | |
| HVP-Net2026.01 | 0.571 | 0.831 | — | — | — | 0.87 | — | — | — | — | |
| InternVideo2V-Enc Params=1.0B, Video Data=102M, Resolution=224, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 0.57 | — | — | — | — | — | — | — | — | — | |
| InternVideo2-1BSetting=ZS2026.05 | 0.57 | — | — | — | — | — | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/322024.03 | 0.569 | 0.822 | — | — | — | 0.892 | — | — | — | — | |
| STM#Pairs=5M2026.03 | 0.569 | 0.841 | — | — | — | 0.897 | — | — | — | — | |
| VLAB_GTraining Data (M)=262023.05 | 0.568 | 0.816 | — | — | — | 0.887 | 227.1 | — | — | — | |
| LaViLaBackbone=CLIP-ViP2026.01 | 0.566 | 0.798 | — | — | — | 0.871 | — | — | — | — | |
| HiTeASample=17M, Modality=Vision-only2023.05 | 0.565 | 0.817 | — | — | — | 0.897 | — | — | — | — | |
| HiTeAVideo-language pretraining (PT)=true2023.12 | 0.565 | 0.817 | — | — | — | 0.897 | — | — | — | — | |
| HiTeA#Pairs=17M2023.03 | 0.565 | 0.817 | — | — | — | 0.897 | — | — | — | — | |
| HiTeAModality=T-V, Protocol=Finetuning2024.12 | 0.565 | — | — | — | — | 0.897 | — | — | — | — | |
| HiTeAsetting=finetuning2025.07 | 0.565 | — | — | — | — | — | — | — | — | — | |
| mPLUG-2Training Data (M)=17, Matching Paradigm=vision-text matching2023.05 | 0.564 | 0.791 | — | — | — | 0.852 | 220.7 | — | — | — | |
| mPLUG-2Sample=417M, Modality=Vision-only2023.05 | 0.564 | 0.791 | — | — | — | 0.852 | — | — | — | — | |
| mPLUG-2Modality=T-V, Protocol=Finetuning2024.12 | 0.564 | — | — | — | — | 0.852 | — | — | — | — | |
| mPLUG-2setting=finetuning2025.07 | 0.564 | — | — | — | — | — | — | — | — | — | |
| CLIP-ViP2026.01 | 0.557 | 0.781 | — | — | — | 0.861 | — | — | — | — | |
| Gemini Embedding 2Setting=ZS2026.05 | 0.5556 | — | — | — | — | — | — | — | — | — | |
| VAST-GZero-shot=true2025.03 | 0.555 | — | — | — | — | — | — | — | — | — | |
| CLIPVIPBackbone=CLIP-ViT-B/16, dual-softmax inference=true2024.03 | 0.553 | 0.82 | — | — | — | 0.893 | — | — | — | — | |
| UniAlign*Modality=T-VA, Zero-shot=true, Tuple-level losses=false2026.02 | 0.552 | — | — | — | — | — | — | — | — | — | |
| VLAB_LTraining Data (M)=262023.05 | 0.551 | 0.819 | — | — | — | 0.876 | 224.6 | — | — | — | |
| UMT-B#Pairs=5M2023.03 | 0.548 | 0.83 | — | — | — | 0.89 | — | — | — | — | |
| UMT#Pairs=5M2026.03 | 0.548 | 0.83 | — | — | — | 0.89 | — | — | — | — | |
| VINDLU#Pairs=5M2026.03 | 0.546 | 0.813 | — | — | — | 0.89 | — | — | — | — | |
| GRAMModality=T-VA, Zero-shot=true2024.12 | 0.542 | — | — | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Zero-shot=true2024.12 | 0.542 | — | — | — | — | 0.793 | — | — | — | — | |
| GRAMModality=T-VA, Zero-shot=true2026.02 | 0.542 | — | — | — | — | — | — | — | — | — | |
| GRAMModality=T-V, Zero-shot=true2024.12 | 0.54 | — | — | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-V, Zero-shot=true2024.12 | 0.54 | — | — | — | — | 0.807 | — | — | — | — | |
| InstAPSetting=Zero-shot2026.04 | 0.54 | 0.782 | — | — | — | 0.845 | — | — | — | — | |
| SINGULARITY#PT=17M, #Train Frame=12022.06 | 0.539 | 0.794 | — | — | — | 0.869 | — | — | — | — | |
| SingularitySample=17M, Modality=Vision-only2023.05 | 0.539 | 0.794 | — | — | — | 0.869 | — | — | — | — | |
| CLIPVIPBackbone=CLIP-ViT-B/32, dual-softmax inference=true2024.03 | 0.538 | 0.796 | — | — | — | 0.865 | — | — | — | — |