Video-to-Text Retrieval on MSR-VTT
64.8Recall@1GRAM Model
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GRAM ModelModality=T-VAS2024.12 | 64.8 | — | 91.5 | — | — | — | |
| GRAM ModelModality=T-VAS, Finetuning=true2024.12 | 64.8 | — | — | — | — | — | |
| VideoCoCaModality=T-V, Zero-shot=true2024.12 | 64.7 | — | — | — | — | — | |
| VideoCoCaModality=T-V, Zero-shot=true2026.02 | 64.7 | — | — | — | — | — | |
| VideoCoCaZero-shot=true2025.07 | 64.7 | — | — | — | — | — | |
| VASTsetting=finetuning2025.07 | 64.3 | — | — | — | — | — | |
| GRAMsetting=finetuning2025.07 | 61.8 | — | — | — | — | — | |
| TEFALModality=T-VA, Finetuning=true2024.12 | 61 | — | — | — | — | — | |
| PMRLsetting=finetuning2025.07 | 60.7 | — | — | — | — | — | |
| InternVideo2s2-6BFinetuning=true2024.03 | 60.2 | — | — | — | — | — | |
| GRAM ModelModality=T-VA2024.12 | 59 | — | 89.1 | — | — | — | |
| GRAM ModelModality=T-VA, Finetuning=true2024.12 | 59 | — | — | — | — | — | |
| UMT-LFinetuning=true2024.03 | 58.6 | — | — | — | — | — | |
| UMT-LModality=T-V, Finetuning and evaluation frames=122024.12 | 58.6 | — | — | — | — | — | |
| UMT-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 58.6 | — | — | — | — | — | |
| UMT-Lsetting=finetuning, evaluation_frames=122025.07 | 58.6 | — | — | — | — | — | |
| InternVideoBackbone=ViT-L/142022.12 | 57.9 | — | — | — | — | — | |
| InternVideo-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 57.9 | — | — | — | — | — | |
| InternVideo-Lsetting=finetuning, evaluation_frames=122025.07 | 57.9 | — | — | — | — | — | |
| VASTModality=T-VA2024.12 | 57.6 | — | 87.4 | — | — | — | |
| VASTModality=T-VAS2024.12 | 57.6 | — | 80.2 | — | — | — | |
| VALOR-LModality=T-VAS, Finetuning=true2024.12 | 57.6 | — | — | — | — | — | |
| VASTModality=T-VA, Finetuning=true2024.12 | 57.6 | — | — | — | — | — | |
| VASTModality=T-VAS, Finetuning=true2024.12 | 57.6 | — | — | — | — | — | |
| HiTeAModality=T-V, Finetuning=true2024.12 | 56.5 | — | — | — | — | — | |
| GRAM ModelModality=T-V2024.12 | 56.4 | — | 87.6 | — | — | — | |
| mPLUG-2Modality=T-V, Finetuning=true2024.12 | 56.4 | — | — | — | — | — | |
| GRAM ModelModality=T-V, Finetuning=true2024.12 | 56.4 | — | — | — | — | — | |
| VidVecV-T Pairs=Text-Only2026.02 | 54.9 | — | — | — | — | — | |
| VidVec-OModel Size=7B, Embedding=Optimized2026.02 | 54.9 | 77.5 | 84.1 | — | — | — | |
| UniAlignModality=T-VA, Zero-shot=true, Tuple-level losses=true2026.02 | 54.6 | — | — | — | — | — | |
| CLIP-ViPReranking (Matching Loss)=true2023.10 | 54.2 | 77.2 | 84.8 | 1 | — | — | |
| InternVideo2-6BV-T Pairs=100M2026.02 | 53.7 | — | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 53.7 | 84.2 | 91.5 | — | — | — | |
| UniAlign*Modality=T-VA, Zero-shot=true, Tuple-level losses=false2026.02 | 53.2 | — | — | — | — | — | |
| GRAMModality=T-VAS, Zero-shot=true2024.12 | 52.9 | — | — | — | — | — | |
| GRAM ModelModality=T-VAS, Evaluation Protocol=Zero-shot2024.12 | 52.9 | — | 82.9 | — | — | — | |
| AuroraRank (r)=64, Input=8x224, Pretraining Parameters=129M, Tunable Parameters=0.1M, Training Protocol=frozen backbone2023.05 | 52.4 | 73.9 | 82 | 1 | — | — | |
| PMRLZero-shot=true2025.07 | 52.4 | — | — | — | — | — | |
| EMCL-NetPre-trained weights=CLIP (ViT-B/32), Inverted softmax=true2022.11 | 51.8 | 80.2 | 88 | 1 | — | — | |
| ViCLIPFinetuning=true2024.03 | 51.8 | — | — | — | — | — | |
| ViCLIPModality=T-V2024.12 | 51.8 | — | — | — | — | — | |
| ViCLIPModality=T-V, Finetuning=true2024.12 | 51.8 | — | — | — | — | — | |
| ViCLIPsetting=finetuning2025.07 | 51.8 | — | — | — | — | — | |
| VideoPrism-gV-T Pairs=600M2026.02 | 51.7 | — | — | — | — | — | |
| GRAMZero-shot=true2025.07 | 51.5 | — | — | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 51.5 | 79.9 | 89.8 | — | — | — | |
| LamRAModel Size=7B2026.02 | 50.9 | 72.8 | 80 | — | — | — | |
| InternVideo2-1BSetting=ZS2026.05 | 50.9 | — | — | — | — | — | |
| UATVR + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 50.9 | 77.4 | 90.5 | — | — | — | |
| T-MASSBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 50.9 | 80.2 | 88 | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 50.8 | 82.7 | 90.4 | — | — | — | |
| UniAdapterRank (r)=512, Input=8x224, Pretraining Parameters=129M, Tunable Parameters=18.8M, Training Protocol=frozen backbone2023.05 | 50.6 | 73.4 | 81.6 | 1 | — | — | |
| GRAMModality=T-VA, Zero-shot=true2024.12 | 50.5 | — | — | — | — | — | |
| GRAM ModelModality=T-VA, Evaluation Protocol=Zero-shot2024.12 | 50.5 | — | 82.2 | — | — | — | |
| GRAMModality=T-VA, Zero-shot=true2026.02 | 50.5 | — | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2024.12 | 50.2 | — | — | — | — | — | |
| VideoPrism-bModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 50.2 | — | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2026.02 | 50.2 | — | — | — | — | — | |
| VideoPrism-bZero-shot=true2025.07 | 50.2 | — | — | — | — | — | |
| X-Pool + OursBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 50.2 | 77.4 | 86.3 | — | — | — | |
| PE-coreLSetting=ZS2026.05 | 50.1 | — | — | — | — | — | |
| LoRARank (r)=32, Input=8x224, Pretraining Parameters=129M, Tunable Parameters=10.6M, Training Protocol=frozen backbone2023.05 | 49.9 | 72 | 81.3 | 2 | — | — | |
| PE-Core-GV-T Pairs=22M2026.02 | 49.9 | — | — | — | — | — | |
| PE-coreGSetting=ZS2026.05 | 49.9 | — | — | — | — | — | |
| UniAdapterRank (r)=128, Input=8x224, Pretraining Parameters=129M, Tunable Parameters=4.6M, Training Protocol=frozen backbone2023.05 | 49.7 | 71.9 | 81.5 | 2 | — | — | |
| UATVR + OursBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 49.7 | 75.6 | 86.4 | — | — | — | |
| GRAMModality=T-V, Zero-shot=true2024.12 | 49.5 | — | — | — | — | — | |
| GRAM ModelModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 49.5 | — | 81.7 | — | — | — | |
| T-MASS + OursBackbone=CLIP-ViT-B/32, Text Augmentation=With2026.05 | 49.5 | 78.1 | 87.5 | — | — | — | |
| CAMoE+DSLmodel_base=CLIP, Dual Softmax Loss=true2021.09 | 49.1 | 74.3 | 84.3 | 2 | 9.9 | — | |
| VASTModality=T-VAS, Zero-shot=true2024.12 | 49 | — | — | — | — | — | |
| VASTModality=T-VAS, Evaluation Protocol=Zero-shot2024.12 | 49 | — | 76.2 | — | — | — | |
| X-CLIPBackbone=ViT-B/162022.07 | 48.9 | 76.8 | 84.5 | 2 | 8.1 | — | |
| X-CLIP2022.12 | 48.9 | — | — | — | — | — | |
| X-CLIPSetting=FT2026.05 | 48.9 | — | — | — | — | — | |
| UATVR + OursBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 48.9 | 76.3 | 87.9 | — | — | — | |
| VASTZero-shot=true2025.07 | 48.8 | — | — | — | — | — | |
| B3Model Size=7B2026.02 | 48.7 | 70.5 | 77.9 | — | — | — | |
| NeighborRetrBackbone=CLIP (ViT-B/32)2025.03 | 48.7 | 74.2 | 84.7 | 2 | 8.4 | 207.5 | |
| X-Pool + OursBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 48.7 | 76 | 84.2 | — | — | — | |
| Gemini Embedding 2Setting=ZS2026.05 | 48.3 | — | — | — | — | — | |
| T-MASSBackbone=CLIP-ViT-B/16, Text Augmentation=With2026.05 | 48.3 | 75.6 | 84.9 | — | — | — | |
| UATVRBackbone=CLIP-ViT-B/16, Text Augmentation=Without2026.05 | 48.1 | 76.3 | 85.4 | — | — | — | |
| MMRet-v1.5Model Size=7B2026.02 | 47.9 | 70.9 | 77.7 | — | — | — | |
| OmniVLReranking (Matching Loss)=true2023.10 | 47.8 | 74.2 | 83.8 | — | — | — | |
| UATVR + OursBackbone=CLIP-ViT-B/32, Text Augmentation=With2026.05 | 47.8 | 74 | 83.9 | — | — | — | |
| CLIP-HhikerInput=120x224, Pretraining Parameters=400M, Tunable Parameters=124M, Training Protocol=full fine-tuning2023.05 | 47.7 | 74.1 | 82.9 | — | — | — | |
| Diffusion2025.03 | 47.7 | 73.8 | 84.5 | 2 | 8.8 | 206 | |
| T-MASSBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 47.7 | 78 | 86.3 | — | — | — | |
| InternVideoBackbone=UniformerV2, Video patch dropping=true2023.10 | 47.4 | 73.2 | 82.6 | 2 | — | — | |
| PE-coreBSetting=ZS2026.05 | 47.3 | — | — | — | — | — | |
| UATVRBackbone=CLIP-ViT-B/32, Text Augmentation=Without2026.05 | 46.9 | 73.8 | 83.8 | — | — | — | |
| X-CLIPBackbone=ViT-B/322022.07 | 46.8 | 73.3 | 84 | 2 | 9.1 | — | |
| HiTeAModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 46.8 | — | — | — | — | — | |
| HBI2025.03 | 46.8 | 74.3 | 84.3 | 2 | 8.9 | 205.4 | |
| DiCoSAQB-Norm=true2025.03 | 46.7 | 75.2 | 84.3 | 2 | 8.9 | 206.2 | |
| TS2Net2022.12 | 46.6 | — | — | — | — | — | |
| EMCL-NetPre-trained weights=CLIP (ViT-B/32), Inverted softmax=false2022.11 | 46.5 | 73.5 | 83.5 | 2 | — | — | |
| EMCL-Net2025.03 | 46.5 | 73.5 | 83.5 | 2 | 8.8 | 203.5 |