Text-to-Video Retrieval on MSR-VTT
64.4Recall@1VAST
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VASTsetting=finetuning2025.07 | 64.4 | — | — | — | — | — | — | — | |
| MiCoBackbone=ViT-g2024.06 | 64.3 | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VAS, Protocol=Finetuning2024.12 | 64 | 89.3 | — | — | — | — | — | — | |
| GRAM ModelModality=T-VAS, Finetuning=true2024.12 | 64 | — | — | — | — | — | — | — | |
| Absolute SOTA2024.06 | 62.8 | — | — | — | — | — | — | — | |
| InternVideo2s2-6BFinetuning=true2024.03 | 62.8 | — | — | — | — | — | — | — | |
| PMRLsetting=finetuning2025.07 | 61.2 | — | — | — | — | — | — | — | |
| GRAMsetting=finetuning2025.07 | 60 | — | — | — | — | — | — | — | |
| UMT-LFinetuning=true2024.03 | 58.8 | — | — | — | — | — | — | — | |
| UMT-LModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 58.8 | 87.1 | — | — | — | — | — | — | |
| UMT-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 58.8 | — | — | — | — | — | — | — | |
| UMT-Lsetting=finetuning, evaluation_frames=122025.07 | 58.8 | — | — | — | — | — | — | — | |
| UniAlignModality=T-VA, Zero-shot=true, Tuple-level losses=true2026.02 | 58.7 | — | — | — | — | — | — | — | |
| vid-TLDRModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 58.5 | 86.9 | — | — | — | — | — | — | |
| vid-TLDRModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 58.5 | — | — | — | — | — | — | — | |
| vid-TLDRsetting=finetuning, evaluation_frames=122025.07 | 58.5 | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Protocol=Finetuning2024.12 | 58.4 | 87 | — | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Finetuning=true2024.12 | 58.4 | — | — | — | — | — | — | — | |
| UniAlign*Modality=T-VA, Zero-shot=true, Tuple-level losses=false2026.02 | 57.7 | — | — | — | — | — | — | — | |
| mPLUG2+Quantization=false, Weight precision=32-bit floating point, Text decoder=LLaMa2025.12 | 57.2 | 85.1 | 80.7 | — | — | — | — | — | |
| CLIP-ViP + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 56.8 | 85.9 | 79.4 | — | — | — | — | — | |
| VASTModality=T-VAS, Protocol=Finetuning2024.12 | 56.6 | 79.4 | — | — | — | — | — | — | |
| VASTModality=T-VAS, Finetuning=true2024.12 | 56.6 | — | — | — | — | — | — | — | |
| VidVecV-T Pairs=Text-Only2026.02 | 56.2 | — | — | — | — | — | — | — | |
| InternVideo2_S2-6BSetting=Zero-shot2024.12 | 55.9 | 85.1 | 78.3 | — | — | — | — | — | |
| InternVideo2-6BV-T Pairs=100M2026.02 | 55.9 | — | — | — | — | — | — | — | |
| InternVideo2_s2-6BEvaluation Protocol=Zero-shot (finetuned with text-video data)2025.07 | 55.9 | 85.1 | 78.3 | — | — | — | — | — | |
| InternVideo2-6BFinetuning data modality=text-video data2025.03 | 55.9 | 85.1 | 78.3 | — | — | — | — | — | |
| VASTModality=T-VA, Protocol=Finetuning2024.12 | 55.8 | 85.9 | — | — | — | — | — | — | |
| VASTModality=T-VA, Finetuning=true2024.12 | 55.8 | — | — | — | — | — | — | — | |
| mPLUG2+Quantization=true, Weight precision=8-bit integer, Text decoder=LLaMa2025.12 | 55.8 | 82.9 | 78.4 | — | — | — | — | — | |
| GRAM ModelModality=T-V, Protocol=Finetuning2024.12 | 55.7 | 86.4 | — | — | — | — | — | — | |
| GRAM ModelModality=T-V, Finetuning=true2024.12 | 55.7 | — | — | — | — | — | — | — | |
| InternVideo-LModality=T-V, Protocol=Finetuning, Number of frames=122024.12 | 55.2 | — | — | — | — | — | — | — | |
| InternVideo-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 55.2 | — | — | — | — | — | — | — | |
| InternVideo-Lsetting=finetuning, evaluation_frames=122025.07 | 55.2 | — | — | — | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 54.9 | 86.8 | 82.6 | — | — | — | — | — | |
| GRAMModality=T-VAS, Zero-shot=true2024.12 | 54.8 | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VAS, Zero-shot=true2024.12 | 54.8 | 82.9 | — | — | — | — | — | — | |
| PMRLZero-shot=true2025.07 | 54.5 | — | — | — | — | — | — | — | |
| VALOR-LModality=T-VAS, Finetuning=true2024.12 | 54.4 | — | — | — | — | — | — | — | |
| VALOR-Lsetting=finetuning2025.07 | 54.4 | — | — | — | — | — | — | — | |
| CLIP-ViPVideo-language pretraining (PT)=true2023.12 | 54.2 | 84.8 | 77.2 | 1 | — | — | — | — | |
| GRAMModality=T-VA, Zero-shot=true2024.12 | 54.2 | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Zero-shot=true2024.12 | 54.2 | 83.9 | — | — | — | — | — | — | |
| GRAMModality=T-VA, Zero-shot=true2026.02 | 54.2 | — | — | — | — | — | — | — | |
| CLIP-ViPBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 54.2 | 84.8 | 77.2 | — | — | — | — | — | |
| Gemini Embedding 2Setting=ZS2026.05 | 53.91 | — | — | — | — | — | — | — | |
| CLIP-ViP + OursBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 53.6 | 84.2 | 77.8 | — | — | — | — | — | |
| RTQVideo-language pretraining (PT)=false2023.12 | 53.4 | 84.4 | 76.1 | 1 | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 53.4 | 86.2 | 81 | — | — | — | — | — | |
| mPLUG-2Modality=T-V, Protocol=Finetuning2024.12 | 53.1 | 84.7 | — | — | — | — | — | — | |
| mPLUG-2Modality=T-V, Finetuning=true2024.12 | 53.1 | — | — | — | — | — | — | — | |
| mPLUG-2setting=finetuning2025.07 | 53.1 | — | — | — | — | — | — | — | |
| GRAMModality=T-V, Zero-shot=true2024.12 | 52.8 | — | — | — | — | — | — | — | |
| GRAM ModelModality=T-V, Zero-shot=true2024.12 | 52.8 | 82.9 | — | — | — | — | — | — | |
| T-MASSModality=T-VA, Protocol=Finetuning2024.12 | 52.7 | 85.6 | — | — | — | — | — | — | |
| T-MASSModality=T-VA, Finetuning=true2024.12 | 52.7 | — | — | — | — | — | — | — | |
| VideoPrism-gV-T Pairs=600M2026.02 | 52.7 | — | — | — | — | — | — | — | |
| T-MASSsetting=finetuning2025.07 | 52.7 | — | — | — | — | — | — | — | |
| T-MASSBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 52.7 | 85.6 | 77.1 | — | — | — | — | — | |
| mPLUG2Quantization=false, Weight precision=32-bit floating point, Text decoder=BERT2025.12 | 52.6 | 83.1 | 75.8 | — | — | — | — | — | |
| ViCLIPFinetuning=true2024.03 | 52.5 | — | — | — | — | — | — | — | |
| ViCLIPModality=T-V, Protocol=Finetuning2024.12 | 52.5 | — | — | — | — | — | — | — | |
| ViCLIPModality=T-V, Finetuning=true2024.12 | 52.5 | — | — | — | — | — | — | — | |
| ViCLIPsetting=finetuning2025.07 | 52.5 | — | — | — | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 52.3 | 87.6 | 77.9 | — | — | — | — | — | |
| TEFALModality=T-VA, Protocol=Finetuning2024.12 | 52 | 86.1 | — | — | — | — | — | — | |
| TEFALModality=T-VA, Finetuning=true2024.12 | 52 | — | — | — | — | — | — | — | |
| TEFALsetting=finetuning2025.07 | 52 | — | — | — | — | — | — | — | |
| InternVideo2-1BSetting=ZS2026.05 | 51.9 | — | — | — | — | — | — | — | |
| CLIP-ViP + OursBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 51.7 | 85.9 | 75.3 | — | — | — | — | — | |
| EMCL-NetPre-trained weights=CLIP (ViT-B/32), inverted softmax=true2022.11 | 51.6 | 85.3 | 78.1 | 1 | — | — | — | — | |
| PE-CoreArchitecture=ViT-G, Zero-shot=true2025.12 | 51.6 | — | — | — | — | — | — | — | |
| GRAMZero-shot=true2025.07 | 51.5 | — | — | — | — | — | — | — | |
| Cap4VideoVideo-language pretraining (PT)=false2023.12 | 51.4 | 83.9 | 75.7 | 1 | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2024.12 | 51.4 | — | — | — | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2024.12 | 51.4 | — | — | — | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2026.02 | 51.4 | — | — | — | — | — | — | — | |
| VideoPrism-bZero-shot=true2025.07 | 51.4 | — | — | — | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/32, Text Augmentation=With2026.05 | 51.4 | 81.3 | 70.2 | — | — | — | — | — | |
| PE-Core-GV-T Pairs=22M2026.02 | 51.2 | — | — | — | — | — | — | — | |
| PE-coreGSetting=ZS2026.05 | 51.2 | — | — | — | — | — | — | — | |
| CLIP-ViPBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 51.2 | 80.4 | 73.9 | — | — | — | — | — | |
| VASTModality=T-VAS, Zero-shot=true2024.12 | 50.7 | — | — | — | — | — | — | — | |
| VASTModality=T-VAS, Zero-shot=true2024.12 | 50.7 | 74.4 | — | — | — | — | — | — | |
| X-Pool + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 50.7 | 85.2 | 76.2 | — | — | — | — | — | |
| VASTZero-shot=true2025.07 | 50.5 | — | — | — | — | — | — | — | |
| CLIP-ViP + OursBackbone=CLIP-ViT-B/32, Text Augmentation=With2026.05 | 50.4 | 84.5 | 73.6 | — | — | — | — | — | |
| PE-coreLSetting=ZS2026.05 | 50.3 | — | — | — | — | — | — | — | |
| T-MASSBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 50.2 | 85.1 | 75.3 | — | — | — | — | — | |
| STOA-VLPVideo-language pretraining (PT)=true2023.12 | 50.1 | 83.8 | 75.5 | — | — | — | — | — | |
| CLIP-ViPBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 50.1 | 84.6 | 74.8 | — | — | — | — | — | |
| STANVideo-language pretraining (PT)=false2023.12 | 50 | 84.1 | 75.2 | 1.5 | — | — | — | — | |
| NeighborRetrBackbone=CLIP (ViT-B/32)2025.03 | 49.5 | 84.1 | 74.1 | 2 | 12.8 | 207.7 | — | — | |
| X-CLIPBackbone=ViT-B/162022.07 | 49.3 | 84.8 | 75.8 | 2 | 12.2 | — | — | — | |
| Zero-shot SOTAmodality=V2023.11 | 49.3 | 73.9 | 68.3 | — | — | — | — | — | |
| VASTModality=T-VA, Zero-shot=true2024.12 | 49.3 | — | — | — | — | — | — | — | |
| VASTModality=T-VA, Zero-shot=true2024.12 | 49.3 | 80 | — | — | — | — | — | — | |
| VASTModality=T-VA, Zero-shot=true2026.02 | 49.3 | — | — | — | — | — | — | — |