Text-to-video retrieval on MSRVTT
61R@1SimVTP
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SimVTPVis Enc. Init=CLIP, Pre-trained Data=WebVid2M, #pairs=2.5M2022.12 | 61 | 87.7 | 94.5 | 1 | |
| UMT-L#Pairs=25M2024.03 | 58.8 | — | — | — | |
| UMT-L + vid-TLDR#Pairs=25M2024.03 | 58.1 | — | — | — | |
| InternVideo#Pairs=646M2024.03 | 55.2 | — | — | — | |
| CLIP-ViP#Pairs=500M2024.03 | 54.2 | — | — | — | |
| SimVTPVis Enc. Init=Kinetics, Pre-trained Data=WebVid2M, #pairs=2.5M2022.12 | 53.6 | 81.9 | 90.7 | 1 | |
| UMT-B#Pairs=25M2024.03 | 51 | — | — | — | |
| UMT-B + vid-TLDR#Pairs=25M2024.03 | 50.9 | — | — | — | |
| DiffusionRet2023.03 | 49 | 75.2 | 82.7 | 2 | |
| OmniVL#Pairs=17M2024.03 | 47.8 | — | — | — | |
| EMCL-Net2023.03 | 47 | 72.3 | 82.6 | 2 | |
| VINDLU#Pairs=25M2024.03 | 46.5 | — | — | — | |
| LocVTPVis Enc. Init=CLIP, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 46.3 | 72.8 | 82 | 2 | |
| BFormerVis Enc. Init=CLIP, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 44.9 | 71.9 | 80 | 2 | |
| Clip4clipVis Enc. Init=CLIP, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 44.5 | 71.4 | 81.6 | 2 | |
| CLIP4Clip#Pairs=400M2024.03 | 44.5 | — | — | — | |
| CLIP4Clip2023.03 | 43.8 | 70.6 | 81.4 | 2 | |
| SCLEvaluation protocol=Fine-tune2022.11 | 43.2 | 76 | 86.7 | — | |
| Singularity#Pairs=17M2024.03 | 42.7 | — | — | — | |
| Clip4ClipPre-training Data Size=400M2022.09 | 42.1 | 71.9 | 81.4 | — | |
| OATransVis Enc. Init=CLIP, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 40.9 | 70.4 | 80.3 | 2 | |
| LAVENDER#Pairs=30M2024.03 | 40.7 | — | — | — | |
| CloverEvaluation protocol=Fine-tune2022.11 | 38.6 | 67.4 | 76.4 | — | |
| All-in-onePre-training Data Size=138M2022.09 | 37.9 | 68.1 | 77.1 | — | |
| All-in-one#Pairs=138M2024.03 | 37.9 | — | — | — | |
| MILESEvaluation protocol=Fine-tune2022.11 | 37.7 | 63.6 | 73.8 | — | |
| BFormerVis Enc. Init=ImageNet, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 37.6 | 64.8 | 75.1 | 3 | |
| MCQEvaluation protocol=Fine-tune2022.11 | 37.6 | 64.8 | 75.1 | — | |
| BridgeFormerPre-training Data Size=5M2022.09 | 37.6 | 64.8 | 75.1 | — | |
| VIOLETv2Pre-training Data Size=5M2022.09 | 37.2 | 64.8 | 75.8 | — | |
| LocVTPVis Enc. Init=ImageNet, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 36.5 | 64.3 | 76.8 | 3 | |
| RegionLearnerVis Enc. Init=ImageNet, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 36.3 | 63.9 | 72.5 | 3 | |
| OATransVis Enc. Init=ImageNet, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 35.8 | 63.4 | 76.5 | 3 | |
| VIOLETEvaluation protocol=Fine-tune2022.11 | 34.5 | 63 | 73.4 | — | |
| VIOLETPre-training Data Size=186M2022.09 | 34.5 | 63 | 73.4 | — | |
| VIOLET#Pairs=138M2024.03 | 34.5 | — | — | — | |
| HiTeA# PT Data=17M, Evaluation mode=Zero-shot2022.12 | 34.4 | 60 | 69.9 | — | |
| ALPROEvaluation protocol=Fine-tune2022.11 | 33.9 | 60.7 | 73.2 | — | |
| ALPROPre-training Data Size=5M2022.09 | 33.9 | 60.7 | 73.2 | — | |
| Clip4Clip# PT Data=400M, Evaluation mode=Zero-shot2022.12 | 31.2 | 53.7 | 64.2 | — | |
| FrozenVis Enc. Init=ImageNet, Pre-trained Data=WebVid2M+CC3M, #pairs=5.5M2022.12 | 31 | 59.5 | 70.5 | 3 | |
| FrozenEvaluation protocol=Fine-tune2022.11 | 31 | 59.5 | 70.5 | — | |
| FrozenPre-training Data Size=5M2022.09 | 31 | 59.5 | 70.5 | — | |
| Frozen#Pairs=5M2024.03 | 31 | — | — | — | |
| VideoClipVis Enc. Init=CLIP, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 30.9 | 55.4 | 66.8 | — | |
| VideoCLIPEvaluation protocol=Fine-tune2022.11 | 30.9 | 55.4 | 66.8 | — | |
| SCLEvaluation protocol=Zero-shot2022.11 | 30.9 | 54.4 | 65 | — | |
| HITVis Enc. Init=Multi-modal, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 30.7 | 60.9 | 73.2 | 2.6 | |
| SupportSetVis Enc. Init=IG65M, ImageNet, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 30.1 | 58.5 | 69.3 | 3 | |
| HiTeA# PT Data=5M, Evaluation mode=Zero-shot2022.12 | 29.9 | 54.2 | 62.9 | — | |
| Singularity# PT Data=5M, Evaluation mode=Zero-shot2022.12 | 28.4 | 50.2 | 59.5 | — | |
| AVLnetVis Enc. Init=ImageNet, Kinetics, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 27.1 | 55.6 | 66.6 | 4 | |
| MMTVis Enc. Init=Multi-modal, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 26.6 | 57.1 | 69.6 | 4 | |
| MILESEvaluation protocol=Zero-shot2022.11 | 26.1 | 47.2 | 56.9 | — | |
| BridgeFormer# PT Data=5M, Evaluation mode=Zero-shot2022.12 | 26 | 46.4 | 56.4 | — | |
| MCQEvaluation protocol=Zero-shot2022.11 | 26 | 46.4 | 56.4 | — | |
| VIOLET# PT Data=183M, Evaluation mode=Zero-shot2022.12 | 25.9 | 49.5 | 59.7 | — | |
| VIOLETEvaluation protocol=Zero-shot2022.11 | 25.9 | 49.5 | 59.7 | — | |
| CloverEvaluation protocol=Zero-shot2022.11 | 25.8 | 49.6 | 60.1 | — | |
| ALPRO# PT Data=5M, Evaluation mode=Zero-shot2022.12 | 24.1 | 44.7 | 55.4 | — | |
| ALPROEvaluation protocol=Zero-shot2022.11 | 24.1 | 44.7 | 55.4 | — | |
| ClipBERTPre-trained Data=COCO, VGen, #pairs=5.6M2022.12 | 22 | 46.8 | 59.9 | 6 | |
| ClipBERTPre-training Data Size=0.2M2022.09 | 22 | 46.8 | 59.9 | — | |
| ClipBERT#Pairs=5.4M2024.03 | 22 | — | — | — | |
| UniVLPre-trained Data=HowTo100M, #pairs=136M2022.12 | 21.2 | 49.6 | 63.1 | 6 | |
| CEVis Enc. Init=Multi-modal, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 20.9 | 48.8 | 62.4 | 6 | |
| Frozen# PT Data=5M, Evaluation mode=Zero-shot2022.12 | 18.7 | 39.5 | 51.6 | — | |
| FrozenEvaluation protocol=Zero-shot2022.11 | 18.7 | 39.6 | 51.6 | — | |
| DECEMBERVis Enc. Init=ImageNet, Kinetics, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 17.5 | 44.3 | 58.6 | 9 | |
| NoiseEstiVis Enc. Init=ImageNet, Kinetics, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 17.4 | 41.6 | 53.6 | 8 | |
| HEROVis Enc. Init=ImageNet, Kinetics, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 16.8 | 43.4 | 57.7 | — | |
| HEROPre-training Data Size=136M2022.09 | 16.8 | 43.4 | 57.7 | — | |
| ActBERTVis Enc. Init=VisGenome, Pre-trained Data=HowTo100M, #pairs=136M2022.12 | 16.3 | 42.8 | 56.9 | 10 | |
| VideoCLIP# PT Data=138M, Evaluation mode=Zero-shot2022.12 | 10.4 | 22.2 | 30 | — | |
| VideoCLIPEvaluation protocol=Zero-shot2022.11 | 10.4 | 22.2 | 30 | — |