Text-to-Video Retrieval on MSR-VTT 1K videos (test)
75.1Recall@10BridgeFormer
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BridgeFormerYear=2021, Video Encoder Input=Raw Videos, PT Dataset=CC3M, WebVid-2M, #Pairs PT=5.5M, evaluation_protocol=fine-tuning2022.01 | 75.1 | 37.6 | 64.8 | 3 | |
| TACOYear=2021, Video Encoder Input=I3D, S3D, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 71.2 | 28.4 | 57.8 | 4 | |
| FrozenYear=2021, Video Encoder Input=Raw Videos, PT Dataset=CC3M, WebVid-2M, #Pairs PT=5.5M, evaluation_protocol=fine-tuning2022.01 | 70.5 | 31 | 59.5 | 3 | |
| MMTYear=2020, Video Encoder Input=S3D, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 69.6 | 26.6 | 57.1 | 4 | |
| SupportSetYear=2021, Video Encoder Input=R(2+1)D-34, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 69.3 | 30.1 | 58.5 | 3 | |
| VLMYear=2021, Video Encoder Input=S3D, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=fine-tuning2022.01 | 67.4 | 28.1 | 55.5 | 4 | |
| VideoCLIPYear=2021, Video Encoder Input=S3D, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=fine-tuning2022.01 | 66.8 | 30.9 | 55.4 | — | |
| AVLnetYear=2021, Video Encoder Input=ResNeXt-101, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 66.6 | 27.1 | 55.6 | 4 | |
| UniVLYear=2020, Video Encoder Input=S3D, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=fine-tuning2022.01 | 63.1 | 21.2 | 49.6 | 6 | |
| ClipBertYear=2021, Video Encoder Input=Raw Videos, PT Dataset=COCO, VisGenome, #Pairs PT=5.6M, evaluation_protocol=fine-tuning2022.01 | 59.9 | 22 | 46.8 | 6 | |
| HEROYear=2021, Video Encoder Input=SlowFast, PT Dataset=TV and HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 57.7 | 16.8 | 43.4 | — | |
| ActBERTYear=2020, Video Encoder Input=ResNet-3D, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=fine-tuning2022.01 | 56.9 | 16.3 | 42.8 | 10 | |
| BridgeFormerYear=2021, Video Encoder Input=Raw Videos, PT Dataset=CC3M, WebVid-2M, #Pairs PT=5.5M, evaluation_protocol=zero-shot2022.01 | 56.4 | 26 | 46.4 | 7 | |
| NoiseEstYear=2021, Video Encoder Input=ResNeXt-101, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=fine-tuning2022.01 | 53.6 | 17.4 | 41.6 | 8 | |
| FrozenYear=2021, Video Encoder Input=Raw Videos, PT Dataset=CC3M, WebVid-2M, #Pairs PT=5.5M, evaluation_protocol=zero-shot2022.01 | 51.6 | 18.7 | 39.5 | 10 | |
| AVLnetYear=2021, Video Encoder Input=ResNeXt-101, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 50.7 | 19.6 | 40.8 | 9 | |
| SupportSetYear=2021, Video Encoder Input=R(2+1)D-34, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 36.2 | 12.7 | 27.5 | 24 | |
| MCNYear=2021, Video Encoder Input=ResNeXt-101, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 33.8 | 10.5 | 25.2 | — | |
| TACOYear=2021, Video Encoder Input=I3D, S3D, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 33.4 | 9.8 | 25 | 29 | |
| ActBERTYear=2020, Video Encoder Input=ResNet-3D, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 33.1 | 8.6 | 23.4 | 36 | |
| MIL-NCEYear=2020, Video Encoder Input=Raw Videos, PT Dataset=HowTo100M, #Pairs PT=120M, evaluation_protocol=zero-shot2022.01 | 32.4 | 9.9 | 24 | 29.6 | |
| MMVYear=2020, Video Encoder Input=Raw Videos, PT Dataset=HowTo100M, AudioSet, #Pairs PT=138M, evaluation_protocol=zero-shot2022.01 | 31.1 | 9.3 | 23 | 38 | |
| VideoCLIPYear=2021, Video Encoder Input=S3D, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=zero-shot2022.01 | 30 | 10.4 | 22.2 | — | |
| VATTYear=2021, Video Encoder Input=Raw Videos, PT Dataset=HowTo100M, AudioSet, #Pairs PT=138M, evaluation_protocol=zero-shot2022.01 | 29.7 | — | — | 49 | |
| NoiseEstYear=2021, Video Encoder Input=ResNeXt-101, PT Dataset=HowTo100M, #Pairs PT=110M, evaluation_protocol=zero-shot2022.01 | 29.3 | 8 | 21.3 | 33 |