Text-to-Video Retrieval on ActivityNet (test)
79.2R@1VidVec
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VidVec2026.02 | 79.2 | 93.4 | 95.2 | — | — | — | |
| InternVideo2-6BTraining Protocol=Fine-tuned on ActivityNet2026.02 | 74.1 | — | — | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=true2023.04 | 70.1 | 90.8 | 95.3 | — | — | — | |
| COSAExample=415M2023.06 | 67.3 | 89 | 95 | — | — | — | |
| UMT-LExample=425M2023.06 | 66.8 | 89.1 | 94.9 | — | — | — | |
| COSA-LExample=417M2023.06 | 66.8 | 87.6 | 93.9 | — | — | — | |
| UMT-LTraining Protocol=Fine-tuned on ActivityNet2026.02 | 66.8 | — | — | — | — | — | |
| VALOR-LExample=433.5M2023.06 | 63.4 | 87.8 | 94.1 | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=false2023.04 | 63.4 | 87.8 | 94.1 | — | — | — | |
| InternVideo2-6BModel Category=Video Foundation Models2026.02 | 63.2 | 85.6 | 92.5 | — | — | — | |
| InternVideo#Example=147.6M, Mod=V, Dual Softmax (+DSL)=true2023.04 | 62.2 | — | — | — | — | — | |
| CLIP-VIP#Example=100M, Mod=V, Dual Softmax (+DSL)=true2023.04 | 61.4 | 85.7 | 92.6 | — | — | — | |
| COSA-BExample=17M2023.06 | 59.3 | 83.8 | 90.9 | — | — | — | |
| LamRAModel Category=Multimodal LLM Embedders2026.02 | 58.5 | 85.4 | 91.6 | — | — | — | |
| HunYuan_tvrMod=V, Dual Softmax (+DSL)=true2023.04 | 57.3 | 84.8 | 93.1 | — | — | — | |
| COSA-BExample=5M2023.06 | 55.6 | 80.7 | 89.1 | — | — | — | |
| PE-Core-GModel Category=Video Foundation Models2026.02 | 54.7 | — | — | — | — | — | |
| VINDLU-BExample=17M2023.06 | 54.4 | 80.7 | 89 | — | — | — | |
| VindLUMethod Category=Non-CLIP Methods, Pretraining Strategy=Pretrained on large-scale video datasets, Retrieval Strategy=Two-stage re-ranking2023.09 | 54.4 | 80.7 | — | — | — | — | |
| CLIP-VIPExample=500M2023.06 | 53.4 | 81.4 | 90 | — | — | — | |
| VideoPrism-gModel Category=Video Foundation Models2026.02 | 52.7 | 79.4 | — | — | — | — | |
| MMRet-v1.5Model Category=Multimodal LLM Embedders2026.02 | 50.6 | 76.6 | 85.8 | — | — | — | |
| VALOR_B#Example=6.5M, Mod=V+A2023.04 | 50.5 | 79.6 | 89.1 | — | — | — | |
| ViCLIPTraining Protocol=Fine-tuned on ActivityNet2026.02 | 49.8 | — | — | — | — | — | |
| HiTeAExample=17M2023.06 | 49.7 | 77.1 | 86.7 | — | — | — | |
| SingularityMethod Category=Non-CLIP Methods, Pretraining Strategy=Pretrained on large-scale video datasets, Retrieval Strategy=Two-stage re-ranking2023.09 | 48.9 | 77 | — | — | — | — | |
| RAPType=Adapter, DSL post-processing=true2024.05 | 48.4 | 76.2 | 86.4 | — | — | 7 | |
| SINGULARITYExample=17M2023.06 | 47.1 | 75.5 | 85.5 | — | — | — | |
| SINGULARITY#Example=17M, Mod=V2023.04 | 47.1 | 75.5 | 85.5 | — | — | — | |
| X-CLIPBackbone=ViT-B/162022.07 | 46.2 | 75.5 | — | — | — | 6.8 | |
| X-CLIPExample=400M2023.06 | 46.2 | 75.5 | — | — | — | — | |
| X-CLIPMod=V2023.04 | 46.2 | 75.5 | — | — | — | — | |
| DCRMod=V2023.04 | 46.2 | 77.3 | 88.2 | — | — | — | |
| UCOFIAMethod Category=CLIP-based Methods2023.09 | 45.7 | 76 | — | 6.6 | — | — | |
| ECLIPSEMod=V+A2023.04 | 45.3 | 75.7 | 86.2 | — | — | — | |
| CLIP4ClipBackbone=ViT-B/16, Aggregation Strategy=seqTransf2022.07 | 44.5 | 75.2 | — | — | — | 6.4 | |
| X-CLIPBackbone=ViT-B/322022.07 | 44.3 | 74.1 | — | — | — | 7.9 | |
| X-CLIPMethod Category=CLIP-based Methods2023.09 | 44.3 | 74.1 | — | 7.9 | — | — | |
| CLIP4ClipBackbone=ViT-B/16, Aggregation Strategy=MeanP2022.07 | 44 | 73.9 | — | — | — | 7 | |
| CLIP4ClipMod=V2023.04 | 43.4 | 70.2 | 80.6 | — | — | — | |
| UMT-LModel Category=Video Foundation Models2026.02 | 42.8 | 69.6 | 79.8 | — | — | — | |
| DiscoVLABackbone=CLIP (ViT-B/32), # Params (M)=0.562025.06 | 41.2 | 72.4 | 83.6 | — | 197.2 | 7.8 | |
| TS2-Net2022.07 | 41 | 73.6 | 84.5 | 2 | 199.1 | — | |
| TS2-Netinverted softmax=false2022.07 | 41 | 73.6 | 84.5 | 2 | — | 8.4 | |
| TS2-NetMethod Category=CLIP-based Methods2023.09 | 41 | 73.6 | — | 8.4 | — | — | |
| VideoMambavisual_backbone=VM, pretraining_pairs=25M, zero-shot=true2024.03 | 41 | 67.5 | 77.8 | — | — | — | |
| TS2-NetMod=V2023.04 | 41 | 73.6 | 84.5 | — | — | — | |
| LanguageBindModel Category=Video Foundation Models2026.02 | 41 | 68.4 | 80 | — | — | — | |
| RAPType=Adapter, DSL post-processing=false2024.05 | 40.8 | 71 | 82.2 | — | — | 8.3 | |
| RAPBackbone=CLIP (ViT-B/32), # Params (M)=1.062025.06 | 40.8 | 71 | 82.2 | — | 194 | 8.3 | |
| CLIP4Clip2022.07 | 40.5 | 73.4 | — | 2 | — | — | |
| CLIP4Clipinverted softmax=false2022.07 | 40.5 | 73.4 | — | 2 | — | 7.5 | |
| CLIP4ClipBackbone=ViT-B/32, Aggregation Strategy=MeanP2022.07 | 40.5 | 72.4 | — | — | — | 7.4 | |
| CLIP4ClipBackbone=ViT-B/32, Aggregation Strategy=seqTransf2022.07 | 40.5 | 72.4 | — | — | — | 7.5 | |
| CLIP4ClipExample=400M2023.06 | 40.5 | 72.4 | — | — | — | — | |
| CLIP4ClipMethod Category=CLIP-based Methods2023.09 | 40.5 | 73.4 | — | 10 | — | — | |
| CLIP4Clip-meanPBackbone=CLIP (ViT-B/32), # Params (M)=123.542025.06 | 40.5 | 72.4 | — | — | — | 7.4 | |
| CLIP4ClipTraining Protocol=Fine-tuned on ActivityNet2026.02 | 40.3 | — | — | — | — | — | |
| VideoMambavisual_backbone=VM, pretraining_pairs=17M, zero-shot=true2024.03 | 40.1 | 65.7 | 76.1 | — | — | — | |
| CLIP4ClipType=Fine-tune, Frozen visual encoder=false2024.05 | 39.4 | 71.1 | 83.3 | — | — | 7.9 | |
| DGLBackbone=CLIP (ViT-B/32), # Params (M)=0.832025.06 | 38.6 | 69.2 | 81.6 | — | 189.4 | 9 | |
| VALOR_E#Example=5.5M, Mod=V2023.04 | 37.5 | 67.9 | 80.4 | — | — | — | |
| VoPF+PBackbone=CLIP (ViT-B/32), # Params (M)=0.42025.06 | 36.1 | 65.5 | 78.5 | — | 180.1 | 10.9 | |
| VideoMambavisual_backbone=VM, pretraining_pairs=5M, zero-shot=true2024.03 | 35.9 | 61.1 | 72.3 | — | — | — | |
| UMTvisual_backbone=ViT, pretraining_pairs=25M, zero-shot=true2024.03 | 35.5 | 60.6 | 71.8 | — | — | — | |
| VLM2Vec-V2Model Category=Multimodal LLM Embedders2026.02 | 35.5 | 54.4 | 61.4 | — | — | — | |
| LF-VILAExample=8.5M2023.06 | 35.3 | 65.4 | — | — | — | — | |
| VoPF+CBackbone=CLIP (ViT-B/32), # Params (M)=14.102025.06 | 35.1 | 63.7 | 77.6 | — | 176.4 | 11.4 | |
| UMTvisual_backbone=ViT, pretraining_pairs=17M, zero-shot=true2024.03 | 33.8 | 59.1 | 70.4 | — | — | — | |
| B3Model Category=Multimodal LLM Embedders2026.02 | 33.8 | 51.9 | 59.6 | — | — | — | |
| SSFType=Adapter2024.05 | 33.2 | 63.6 | 77 | — | — | 11.3 | |
| VoPPType=Prompt2024.05 | 32.8 | 62.3 | 75.4 | — | — | 12.3 | |
| VLM2VecModel Category=Multimodal LLM Embedders2026.02 | 32.8 | 52.9 | 61.3 | — | — | — | |
| VoPCType=Prompt2024.05 | 32.6 | 62.5 | 76.5 | — | — | 12 | |
| UniME-V2Model Category=Multimodal LLM Embedders2026.02 | 31.4 | 50.9 | 59.9 | — | — | — | |
| Singularityvisual_backbone=Swin, pretraining_pairs=5M, zero-shot=true2024.03 | 30.8 | 55.9 | 66.3 | — | — | — | |
| InternVideovisual_backbone=ViT, pretraining_pairs=640M, zero-shot=true2024.03 | 30.7 | 57.4 | 70.2 | — | — | — | |
| InternVideo-LModel Category=Video Foundation Models2026.02 | 30.7 | — | — | — | — | — | |
| Singularityvisual_backbone=Swin, pretraining_pairs=17M, zero-shot=true2024.03 | 30.6 | 55.6 | 66.9 | — | — | — | |
| TACO#Example=136M, Mod=V2023.04 | 30.4 | 61.2 | — | — | — | — | |
| ST-AdapterType=Adapter2024.05 | 29.8 | 59.5 | 73.7 | — | — | 14.5 | |
| HiT2022.07 | 29.6 | 60.7 | — | — | — | — | |
| SSB2022.07 | 29.2 | 61.6 | — | — | — | — | |
| Support SetMethod Category=Non-CLIP Methods2023.09 | 29.2 | 61.6 | — | — | — | — | |
| Support-set#Example=136M, Mod=V2023.04 | 29.2 | 61.6 | — | — | — | — | |
| CoOpType=Prompt2024.05 | 29.1 | 57.3 | 72.2 | — | — | 14.2 | |
| Gabeur et al.#Example=136M, Mod=V+A+S2023.04 | 29 | 61.7 | — | — | — | — | |
| MMTinverted softmax=false, pretrained=true2022.07 | 28.7 | 61.4 | — | 3.3 | — | 16 | |
| MMT2022.07 | 28.7 | 61.4 | — | — | — | 16 | |
| MMT#Example=136M, Mod=V+A2023.04 | 28.7 | 61.4 | — | — | — | — | |
| UMTvisual_backbone=ViT, pretraining_pairs=5M, zero-shot=true2024.03 | 28.3 | 53 | 64.2 | — | — | — | |
| VPTType=Prompt2024.05 | 27.8 | 56 | 70 | — | — | 20.2 | |
| LoRAType=Adapter2024.05 | 27.7 | 55.8 | 69.3 | — | — | 18.8 | |
| MMTMethod Category=Non-CLIP Methods2023.09 | 26.6 | 57.1 | — | 16 | — | — | |
| UNITEModel Category=Multimodal LLM Embedders2026.02 | 25.8 | 42.9 | 51.6 | — | — | — | |
| TT-CE+2022.07 | 23.5 | 57.2 | — | — | — | — | |
| All-in-oneMethod Category=Non-CLIP Methods2023.09 | 22.4 | 53.7 | — | — | — | — | |
| All-in-one#Example=138M, Mod=V2023.04 | 22.4 | 53.7 | 67.7 | — | — | — | |
| CLIP4ClipType=Fine-tune, Frozen visual encoder=true2024.05 | 21.6 | 46.5 | 60.3 | — | — | 37.6 | |
| ClipBERTinverted softmax=false2022.07 | 21.3 | 49 | 63.5 | 6 | — | — |