Text-to-Video Retrieval on YouCook2 (val)
1,510R@1MIL-NCE
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MIL-NCEBackbone=S3D, Labeled dataset used=None2019.12 | 1,510 | 3,800 | 5,120 | 10 | |
| MIL-NCEPre-training=HowTo100M, Visual Backbone=S3D, Batch Size=81922022.06 | 1,510 | 3,800 | 5,120 | 10 | |
| MMV FACPre-training=HowTo100M+AudioSet, Visual Backbone=TSM-50, Batch Size=40962022.06 | 1,150 | 3,020 | 4,150 | 16 | |
| MIL-NCEBackbone=I3D, Labeled dataset used=None2019.12 | 1,140 | 3,060 | 4,200 | 16 | |
| ActBertPre-training=HowTo100M, Visual Backbone=ResNet-3D2022.06 | 960 | 2,670 | 3,800 | 19 | |
| Miech et al.Labeled dataset used=ImNet + K400 + YouCook22019.12 | 820 | 2,450 | 3,530 | 24 | |
| Miech et al+FTPre-training=HowTo100M, Visual Backbone=ResNeXt-1012022.06 | 820 | 2,450 | 3,530 | 24 | |
| Miech et al.Labeled dataset used=ImNet + K4002019.12 | 610 | 1,730 | 2,480 | 46 | |
| Semantic Role Aware Correlation TransformerPre-training=No, Visual Backbone=ResNeXt-101, Batch Size=322022.06 | 530 | 1,450 | 2,080 | 77 | |
| HGRPre-training=No, Visual Backbone=ResNeXt-101, Batch Size=322022.06 | 470 | 1,410 | 2,000 | 87 | |
| HGLMM FV CCALabeled dataset used=ImNet + K400 + YouCook22019.12 | 460 | 1,430 | 2,160 | 75 | |
| HGLMMPre-training=No2022.06 | 460 | 1,430 | 2,160 | 75 | |
| Miech et alPre-training=No, Visual Backbone=ResNeXt-1012022.06 | 420 | 1,370 | 2,150 | 65 | |
| COOTTrainSet=HowTo100M+YouCook22020.11 | 77.2 | 95.8 | 97.5 | 1 | |
| MIL-NCETrainSet=HowTo100M2020.11 | 61.9 | 89.4 | 98.9 | 1 | |
| Miech et al.TrainSet=HowTo100M+YouCook22020.11 | 59.6 | 86 | 93.6 | 1 | |
| COOTTrainSet=YouCook22020.11 | 50.4 | 79.4 | 87.4 | 1.3 | |
| Miech et al.TrainSet=HowTo100M2020.11 | 43.1 | 68.6 | 79.1 | 2 | |
| Miech et al.TrainSet=YouCook22020.11 | 32.3 | 59.2 | 70.9 | 4 | |
| TAN (S1 + S2)Trained on YC2=false, Zero-shot=true, Training Stage=Stage-1+S2 (initialization followed by co-training)2022.04 | 20.1 | 45.5 | 59.5 | 7 | |
| TaCoTrained on YC2=false, Zero-shot=true2022.04 | 19.9 | 43.2 | 55.7 | 8 | |
| TAN (S1)Trained on YC2=false, Zero-shot=true, Training Stage=Stage-1 (initialization)2022.04 | 16.8 | 41.3 | 54.8 | 8 | |
| COOTTrainSet=HowTo100M+YouCook22020.11 | 16.7 | 40.2 | 52.3 | 9 | |
| COOTVisual Backbone=S3D (HowTo100M), Batch Size=64, Pre-training=true2022.06 | 16.7 | 40.2 | 52.3 | 9 | |
| MMCVVisual Backbone=S3D (Kinetics), Batch Size=32, Pre-training=true2022.06 | 16.6 | 37.4 | 48.3 | 12 | |
| TACoVisual Backbone=S3D (HowTo100M), Batch Size=128, Pre-training=true2022.06 | 16.6 | 40.3 | 53.1 | 9 | |
| MILNCEfv=S3D-G, zero-shot=true2020.06 | 15.1 | 38 | 51.2 | 10 | |
| MIL-NCETrained on YC2=false, Zero-shot=true2022.04 | 15.1 | 38 | 51.2 | 10 | |
| MIL-NCETrainSet=HowTo100M2020.11 | 15.1 | 38 | 51.2 | 10 | |
| MIL-NCE (reproduced)Trained on YC2=false, Zero-shot=true, Source=Reproduced in [83]2022.04 | 13.9 | 36.3 | 48.9 | 11 | |
| HTM-AAVideo-Text Training Data=HTM-AA (auto-aligned ASRs) [17], Evaluation Protocol=Zero-shot2023.10 | 13.4 | 32.2 | 43.5 | 15 | |
| HowToCaptionVideo-Text Training Data=HowToCaption (ours), Evaluation Protocol=Zero-shot2023.10 | 13.4 | 33.1 | 44.1 | 15 | |
| MMV FACfv=TSM-50x2, zero-shot=true2020.06 | 11.7 | 33.4 | 45.4 | 13 | |
| MMV FACfv=TSM-50, zero-shot=true2020.06 | 11.5 | 30.2 | 41.5 | 16 | |
| ActBERTTrained on YC2=false, Zero-shot=true2022.04 | 9.6 | 26.7 | 38 | 19 | |
| ActBERTTrainSet=HowTo100M2020.11 | 9.6 | 26.7 | 38 | 19 | |
| MMV FACfv=S3D-G, zero-shot=true2020.06 | 9 | 25.7 | 37.2 | 20 | |
| HowTo100M (distant supervision)Video-Text Training Data=HowTo100M with distant supervision [31], Evaluation Protocol=Zero-shot2023.10 | 8.3 | 21.5 | 30.3 | 34 | |
| Miech et al.TrainSet=HowTo100M+YouCook22020.11 | 8.2 | 24.5 | 35.3 | 24 | |
| UniVL (FT-Joint)Visual Backbone=S3D (HowTo100M), Batch Size=32, Pre-training=true2022.06 | 7.7 | 23.9 | 34.7 | 21 | |
| WebVid2MVideo-Text Training Data=WebVid2M [4], Evaluation Protocol=Zero-shot2023.10 | 7.3 | 20.7 | 29 | 46 | |
| RoMEVisual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=32, Pre-training=false2022.06 | 6.3 | 16.9 | 25.2 | 53 | |
| HowTo100M (ASRs)Video-Text Training Data=HowTo100M with ASRs [17], Evaluation Protocol=Zero-shot2023.10 | 6.1 | 16.2 | 23.6 | 69 | |
| Miech et al.TrainSet=HowTo100M2020.11 | 6.1 | 17.3 | 24.8 | 46 | |
| COOTTrainSet=YouCook22020.11 | 5.9 | 16.7 | 24.8 | 49.7 | |
| COOTVisual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=32, Pre-training=false2022.06 | 5.9 | 16.7 | 24.8 | 49.7 | |
| VideoCC3MVideo-Text Training Data=VideoCC3M [39], Evaluation Protocol=Zero-shot2023.10 | 5.3 | 15.1 | 21.7 | 84 | |
| Satar et al.Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics) + Faster R-CNN (MS COCO), Batch Size=32, Pre-training=false2022.06 | 5.3 | 14.5 | 20.8 | 77 | |
| TACoVisual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=128, Pre-training=false2022.06 | 4.9 | 14.7 | 21.7 | 63 | |
| RoMEVisual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=322022.06 | 4.9 | 15.9 | 24.1 | 55 | |
| HGR*Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=32, Pre-training=false2022.06 | 4.8 | 14 | 20.3 | 85 | |
| SwAMPVisual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=128, Pre-training=false2022.06 | 4.8 | 14.5 | 22.5 | 57 | |
| HGR*Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics) + Faster R-CNN (MS COCO), Batch Size=32, Pre-training=false2022.06 | 4.7 | 14.1 | 20 | 87 | |
| HGLMMTrainSet=YouCook22020.11 | 4.6 | 14.3 | 21.6 | 75 | |
| HGLMMVisual Backbone=Fisher Vectors, Pre-training=false2022.06 | 4.6 | 14.3 | 21.6 | 75 | |
| Satar et al.Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=32, Pre-training=false2022.06 | 4.5 | 13.2 | 20 | 85 | |
| Miech et al.TrainSet=YouCook22020.11 | 4.2 | 13.7 | 21.5 | 65 | |
| Miech et al.Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Pre-training=false2022.06 | 4.2 | 13.7 | 21.5 | 65 | |
| HGR*Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics) + Faster R-CNN (MS COCO), Batch Size=322022.06 | 4.2 | 13.2 | 18.9 | 84 | |
| HGR*Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=322022.06 | 4.1 | 13 | 19 | 85 | |
| UniVL (v2-FT-Joint)Visual Backbone=Resnet-152 (ImageNet) + ResNeXt-101 (Kinetics), Batch Size=32, Pre-training=false2022.06 | 3.4 | 10.8 | 17.8 | 76 | |
| RandomLabeled dataset used=None2019.12 | 3 | 15 | 30 | 1,675 | |
| RandomPre-training=No2022.06 | 3 | 15 | 30 | 1,675 | |
| RandomTrainSet=-2020.11 | 0.21 | 1.09 | 2.19 | 229 | |
| RandomTrainSet=-2020.11 | 0.03 | 0.15 | 0.3 | 1,675 | |
| RandomPre-training=false2022.06 | 0.03 | 0.15 | 0.3 | 1,675 |