Text-to-Video Retrieval on DiDeMo (test)
70.5R@1COSA
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| COSAExample=415M2023.06 | 70.5 | 89.3 | 92.4 | — | — | — | — | — | — | — | — | |
| UMT-LExample=425M2023.06 | 70.4 | 90.1 | 93.5 | — | — | — | — | — | — | — | — | |
| COSA-LExample=417M2023.06 | 68.3 | 88.1 | 91.7 | — | — | — | — | — | — | — | — | |
| TESTAPre-training Data=5M, GFLOPS=786, Evaluation Protocol=Zero-shot, Frames=322023.10 | 64.9 | 88.7 | 91.8 | — | — | — | — | — | — | — | — | |
| COSA-BExample=17M2023.06 | 64.1 | 86.1 | 90.6 | — | — | — | — | — | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=true2023.04 | 61.5 | 85.3 | 90.4 | — | — | — | — | — | — | — | — | |
| VINDLU#Data=25M, #Frames=4, Time=822022.12 | 61.2 | 85.8 | 91 | — | — | — | 79.3 | — | — | — | — | |
| TESTA#PT Data=5M, #Frame=96, GFLOPS=13812023.10 | 61.2 | 87.2 | 91.5 | — | — | — | — | — | — | — | — | |
| BLIPPre-training Data=129M, GFLOPS=707, Evaluation Protocol=Zero-shot, Frames=322023.10 | 60.9 | 84.9 | 91 | — | — | — | — | — | — | — | — | |
| VINDLU-L#Data=25M, #Frames=4, Time=1782022.12 | 59.8 | 86.6 | 91.5 | — | — | — | 79.3 | — | — | — | — | |
| VINDLU#Data=17M, #Frames=4, Time=382022.12 | 59.2 | 84.1 | 89.5 | — | — | — | 77.6 | — | — | — | — | |
| VINDLU-BExample=17M2023.06 | 59.2 | 84.1 | 89.5 | — | — | — | — | — | — | — | — | |
| VindLUMethod Category=Non-CLIP Methods, Pretraining Strategy=Pretrained on large-scale video datasets, Retrieval Strategy=Two-stage re-ranking2023.09 | 59.2 | 84.1 | — | — | — | — | — | — | — | — | — | |
| VIRTUE v1 7Bretrieval_stage=Dual, model_scale=7B, zero_shot=true2026.01 | 58.8 | 81 | — | — | — | — | — | — | — | — | — | |
| InternVideo#Example=147.6M, Mod=V, Dual Softmax (+DSL)=true2023.04 | 57.9 | — | — | — | — | — | — | — | — | — | — | |
| InternVideo2 6Bretrieval_stage=Dual, model_scale=6B, zero_shot=true2026.01 | 57.9 | 80 | — | — | — | — | — | — | — | — | — | |
| COSA-BExample=5M2023.06 | 57.8 | 80.6 | 87.9 | — | — | — | — | — | — | — | — | |
| TESTA#PT Data=5M, #Frame=32, GFLOPS=4202023.10 | 57.7 | 83.3 | 89.4 | — | — | — | — | — | — | — | — | |
| VALOR-LExample=433.5M2023.06 | 57.6 | 83.3 | 88.8 | — | — | — | — | — | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=false2023.04 | 57.6 | 83.3 | 88.8 | — | — | — | — | — | — | — | — | |
| HiTeA# PT Data=17M, Inference Setting=Fine-tuning2022.12 | 56.5 | 81.7 | 89.7 | — | — | — | — | — | — | — | — | |
| HiTeAExample=17M2023.06 | 56.5 | 81.7 | 89.7 | — | — | — | — | — | — | — | — | |
| MuLTI-L#PT=5.5M, Post-processing=DSL or QB-Norm2023.03 | 56.5 | 80.2 | 87 | — | — | — | — | 73.3 | — | — | — | |
| mPLUG-2Example=417M2023.06 | 56.4 | 79.1 | 85.2 | — | — | — | — | — | — | — | — | |
| VIRTUE v2 7Bretrieval_stage=Dual, model_scale=7B, zero_shot=true2026.01 | 56.4 | 80.4 | — | — | — | — | — | — | — | — | — | |
| VidVec-ZSModel Size=7B, Evaluation Protocol=Zero-shot, Reranking=K=1002026.02 | 55.7 | 80.7 | 85.4 | — | — | — | — | — | — | — | — | |
| VASTSample=443M, Zero-shot=true, Modality=Omni-modal (Vision/Audio/Subtitle)2023.05 | 55.5 | 74.3 | 79.6 | — | — | — | — | — | — | — | — | |
| CLIP-ViPBackbone=ViT-B/16, Post-processing DSL=true2022.09 | 55.3 | 82 | 89.3 | — | 1 | — | — | — | — | — | — | |
| CLIP-VIP#Example=100M, Mod=V, Dual Softmax (+DSL)=true2023.04 | 55.3 | 82 | 89.3 | — | — | — | — | — | — | — | — | |
| VINDLU#Data=5M, #Frames=4, Time=152022.12 | 54.6 | 81.3 | 89 | — | — | — | 75 | — | — | — | — | |
| VINDLU#PT Data=5M, #Frame=4, GFLOPS=932023.10 | 54.6 | 81.3 | 89 | — | — | — | — | — | — | — | — | |
| SINGULARITYExample=17M2023.06 | 53.9 | 79.4 | 86.9 | — | — | — | — | — | — | — | — | |
| SINGULARITY#Example=17M, Mod=V2023.04 | 53.9 | 79.4 | 86.9 | — | — | — | — | — | — | — | — | |
| CLIP-ViPBackbone=ViT-B/32, Post-processing DSL=true2022.09 | 53.8 | 79.6 | 86.5 | — | 1 | — | — | — | — | — | — | |
| VidVec-OModel Size=7B, Optimization Data=60K text-only in-context pairs2026.02 | 53.7 | 79.4 | 85 | — | — | — | — | — | — | — | — | |
| LAVENDER#Data=30M, #Frames=4, Time=6402022.12 | 53.4 | 78.6 | 85.3 | — | — | — | 72.4 | — | — | — | — | |
| LAVENDERExample=30M2023.06 | 53.4 | 78.6 | 85.3 | — | — | — | — | — | — | — | — | |
| LAVENDER#Example=30M, Mod=V2023.04 | 53.4 | 78.6 | 85.3 | — | — | — | — | — | — | — | — | |
| Singularity#Data=17M, #Frames=1-4, Time=292022.12 | 53.1 | 79.9 | 88.1 | — | — | — | 73.7 | — | — | — | — | |
| SingularityMethod Category=Non-CLIP Methods, Pretraining Strategy=Pretrained on large-scale video datasets, Retrieval Strategy=Two-stage re-ranking2023.09 | 53.1 | 79.9 | — | — | — | — | — | — | — | — | — | |
| OmniVL#Data=17M, #Frames=1-8, Time=169*2022.12 | 52.4 | 79.5 | 85.4 | — | — | — | 72.4 | — | — | — | — | |
| OmniVLExample=17M2023.06 | 52.4 | 79.5 | 85.4 | — | — | — | — | — | — | — | — | |
| VALOR_B#Example=6.5M, Mod=V+A2023.04 | 52.2 | 80.8 | 86.8 | — | — | — | — | — | — | — | — | |
| HunYuan_tvrMod=V, Dual Softmax (+DSL)=true2023.04 | 52.1 | 78.2 | 85.7 | — | — | — | — | — | — | — | — | |
| Cap4VideoBackbone=CLIP, Video Length=64 frames, Max Text Words=64, Interaction Layers=12022.12 | 52 | 79.4 | 87.5 | 1 | 10.5 | — | — | — | — | — | — | |
| HiTeA# PT Data=5M, Inference Setting=Fine-tuning2022.12 | 51.8 | 79.1 | 85.3 | — | — | — | — | — | — | — | — | |
| HiTeA#PT Data=5M, #Frame=12, GFLOPS=982023.10 | 51.8 | 79.1 | 85.3 | — | — | — | — | — | — | — | — | |
| CLIP-ViPBackbone=ViT-B/16, Post-processing DSL=false2022.09 | 50.5 | 78.4 | 87.1 | — | 1 | — | — | — | — | — | — | |
| CLIP-ViP#Data=500M, #Frames=1-12, Time=984*2022.12 | 50.5 | 78.4 | 87.1 | — | — | — | 72 | — | — | — | — | |
| CLIP-VIPExample=500M2023.06 | 50.5 | 78.4 | 87.1 | — | — | — | — | — | — | — | — | |
| CLIP-ViP#PT Data=100M, #Frame=12, GFLOPS=2122023.10 | 50.5 | 78.4 | 87.1 | — | — | — | — | — | — | — | — | |
| MuLTI-L#PT=5.5M2023.03 | 50.5 | 78.5 | 86.2 | — | — | — | — | 69.9 | — | — | — | |
| CloverPre-training dataset=W2M+C3M, Evaluation Protocol=Fine-tune2022.07 | 50.1 | 76.7 | 85.6 | 1 | — | — | — | — | — | — | — | |
| Clover#PT=5.5M2023.03 | 50.1 | 76.7 | 85.6 | — | — | — | — | 69 | — | — | — | |
| DRLBackbone=ViT-B/162022.09 | 49 | 76.5 | 84.5 | — | 2 | — | — | — | — | — | — | |
| DRL2022.12 | 49 | 76.5 | 84.5 | 2 | — | — | — | — | — | — | — | |
| DCRMod=V2023.04 | 49 | 76.5 | 84.5 | — | — | — | — | — | — | — | — | |
| DiffusionRetQB-Norm=true2023.03 | 48.9 | 75.5 | 83.3 | 2 | 14.1 | 207.7 | — | — | — | — | — | |
| CLIP-ViPBackbone=ViT-B/32, Post-processing DSL=false2022.09 | 48.6 | 77.1 | 84.4 | — | 2 | — | — | — | — | — | — | |
| UMT-LSample=425M, Zero-shot=true, Modality=Vision-only2023.05 | 48.6 | 72.9 | 79 | — | — | — | — | — | — | — | — | |
| PAU2023.09 | 48.6 | 76 | 84.5 | 2 | 12.9 | — | — | — | — | — | — | |
| MuLTI-B#PT=5.5M, Post-processing=DSL or QB-Norm2023.03 | 48.3 | 75.4 | 83.5 | — | — | — | — | 67.2 | — | — | — | |
| NeighborRetr2025.03 | 48.2 | 76.7 | 84.9 | 2 | 11.9 | — | — | — | — | — | — | |
| DRLBackbone=ViT-B/322022.09 | 47.9 | 73.8 | 82.7 | — | 2 | — | — | — | — | — | — | |
| VIOLETv2Example=5M2023.06 | 47.9 | 76.5 | 84.1 | — | — | — | — | — | — | — | — | |
| MuLTI-S#PT=5.5M, Post-processing=DSL or QB-Norm2023.03 | 47.9 | 73 | 82.6 | — | — | — | — | 66.1 | — | — | — | |
| X-CLIPBackbone=ViT-B/162022.07 | 47.8 | 79.3 | — | — | 12.6 | — | — | — | — | — | — | |
| X-CLIPExample=400M2023.06 | 47.8 | 79.3 | — | — | — | — | — | — | — | — | — | |
| X-CLIP#PT Data=400M, #Frame=64, GFLOPS=10862023.10 | 47.8 | 79.3 | — | — | — | — | — | — | — | — | — | |
| X-CLIPMod=V2023.04 | 47.8 | 79.3 | — | — | — | — | — | — | — | — | — | |
| Singularity# PT Data=5M, Inference Setting=Fine-tuning2022.12 | 47.4 | 75.2 | 84 | — | — | — | — | — | — | — | — | |
| LAVENDER# PT Data=5M, Inference Setting=Fine-tuning2022.12 | 47.4 | 74.7 | 82.4 | — | — | — | — | — | — | — | — | |
| Singularity#PT Data=5M, #Frame=32, GFLOPS=5892023.10 | 47.4 | 75.2 | 84 | — | — | — | — | — | — | — | — | |
| TS2-Net#PT=400M, Post-processing=DSL or QB-Norm2023.03 | 47.4 | 74.1 | 82.4 | — | — | — | — | 66.1 | — | — | — | |
| CLIP4Clip-MeanPBackbone=ViT-B/162022.07 | 47.3 | 75.1 | — | — | 13 | — | — | — | — | — | — | |
| CLIP4Clip-seqTransfBackbone=ViT-B/162022.07 | 47.2 | 73.4 | — | — | 13.5 | — | — | — | — | — | — | |
| RAPType=Adapter, DSL post-processing=true2024.05 | 47.1 | 74.1 | 82.4 | — | 13.9 | — | — | — | — | — | — | |
| BaselineCondition=without PAU2023.09 | 47 | 76 | 86.1 | 2 | 10.7 | — | — | — | — | — | — | |
| HBI2023.03 | 46.9 | 74.9 | 82.7 | 2 | 12.1 | 204.5 | — | — | — | — | — | |
| HBI2025.03 | 46.9 | 74.9 | 82.7 | 2 | 12.1 | — | — | — | — | — | — | |
| DiffusionRet2023.03 | 46.7 | 74.7 | 82.7 | 2 | 14.3 | 204.1 | — | — | — | — | — | |
| Diffusion2025.03 | 46.7 | 74.7 | 82.7 | 2 | 14.3 | — | — | — | — | — | — | |
| VIRTUE-Embed 7Bretrieval_stage=Single, model_category=MLLM-based, model_scale=7B, zero_shot=true2026.01 | 46.6 | 70.8 | — | — | — | — | — | — | — | — | — | |
| UCOFIA2023.09 | 46.5 | 74.8 | 84.4 | 2 | 13.4 | — | — | — | — | — | — | |
| UCOFIAMethod Category=CLIP-based Methods2023.09 | 46.5 | 74.8 | — | — | 13.4 | — | — | — | — | — | — | |
| VoPF+CParams (M)=14.1 (11.785%)2022.11 | 46.4 | 71.9 | 81.5 | 2 | 13.6 | — | — | — | — | — | — | |
| LamRAModel Size=7B, Evaluation Protocol=Zero-shot2026.02 | 46.3 | 71.8 | 80.5 | — | — | — | — | — | — | — | — | |
| LamRAModel Size=7B, Training Scale=Large vision-text scale2026.02 | 46.3 | 71.8 | 80.5 | — | — | — | — | — | — | — | — | |
| mPLUG-2Zero-shot=true, Pre-training Data=17M2023.02 | 45.7 | 71.1 | 79.2 | — | — | — | — | — | — | — | — | |
| DICOSA2023.03 | 45.7 | 74.6 | 83.5 | 2 | 11.7 | 203.8 | — | — | — | — | — | |
| CLIP2TVExtension=SD2021.11 | 45.5 | 69.7 | 80.6 | 2 | 17.1 | — | — | — | — | — | — | |
| CLIP2TVBackbone=ViT-B/322022.09 | 45.5 | 69.7 | 80.6 | — | 2 | — | — | — | — | — | — | |
| CLIP2TV2023.09 | 45.5 | 69.7 | 80.6 | 2 | 17.1 | — | — | — | — | — | — | |
| VoPF+PParams (M)=0.4 (0.328%)2022.11 | 45.3 | 72.3 | 80.4 | 2 | 13.8 | — | — | — | — | — | — | |
| EMCL-Net2023.03 | 45.3 | 74.2 | 82.3 | 2 | 12.3 | 201.8 | — | — | — | — | — | |
| X-CLIPBackbone=ViT-B/322022.07 | 45.2 | 74 | — | — | 14.6 | — | — | — | — | — | — | |
| X-CLIP# PT Data=400M, Inference Setting=Fine-tuning2022.12 | 45.2 | 74 | — | — | — | — | — | — | — | — | — | |
| X-CLIP2023.09 | 45.2 | 74 | — | — | 14.6 | — | — | — | — | — | — | |
| X-CLIPMethod Category=CLIP-based Methods2023.09 | 45.2 | 74 | — | — | 14.6 | — | — | — | — | — | — | |
| MuLTI-B#PT=5.5M2023.03 | 45.2 | 74.6 | 82.2 | — | — | — | — | 65.2 | — | — | — |