Text-to-Video Retrieval on YouCook2
83.7Recall@10MIL-NCE*
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| MIL-NCE*Prediction=Caption Avg.2021.04 | 83.7 | 46.6 | 74.3 | — | — | — | — | |
| MCNPrediction=Caption Avg.2021.04 | 81.4 | 53.4 | 75 | — | — | — | — | |
| VASTModality=Audio/Subtitle2023.05 | 80.8 | 50.4 | 74.3 | — | — | — | — | |
| Miech et al.Prediction=Caption Avg.2021.04 | 79.1 | 43.1 | 68.6 | — | — | — | — | |
| MCNPrediction=MV-Video2021.04 | 78.4 | 38.8 | 67.7 | — | — | — | — | |
| MCNPrediction=MV-Clip2021.04 | 76.8 | 38.8 | 67.4 | — | — | — | — | |
| VideoCLIP2023.03 | 75 | 32.2 | 62.6 | — | — | — | — | |
| UniVL + MELTRregularization=true2023.03 | 74.8 | 33.7 | 63.1 | 3 | — | — | — | |
| MELTRModality=Vision-only2023.05 | 74.8 | 33.7 | 63.1 | — | — | — | — | |
| UniVL + MELTR-regularization=false2023.03 | 73.3 | 33.4 | 62.5 | 3 | — | — | — | |
| TACO2023.03 | 72.7 | 29.6 | 59.7 | 9 | — | — | — | |
| TACoEvaluation Protocol=Finetuned, Video Backbone=S3D-HM2021.08 | 72.7 | 29.6 | 59.7 | — | — | 4 | — | |
| OmniVecModality=Video+Text, Fine-tuned=true2025.07 | 70.8 | — | — | — | — | — | — | |
| UniVL-Aligntraining_loss=LAlign2023.03 | 70 | 28.9 | 57.6 | 4 | — | — | — | |
| UniVLModality=Vision-only2023.05 | 70 | 28.9 | 57.6 | — | — | — | — | |
| UniVLapproach=FT-Align2020.02 | 70 | 28.9 | 57.6 | 4 | — | — | — | |
| UniVL(v3)Evaluation Protocol=Finetuned, Video Backbone=S3D-HM2021.08 | 70 | 28.9 | 57.6 | — | — | 4 | — | |
| OmniVec2Modality=Video+Text, Zero-shot=true2025.07 | 69.9 | — | — | — | — | — | — | |
| VLMModality=Vision-only2023.05 | 69.4 | 27.1 | 56.9 | — | — | — | — | |
| VLMPre-training Paradigm=Multi-task Pre-training2021.05 | 69.38 | 27.05 | 56.88 | — | — | 4 | — | |
| TACoEvaluation Protocol=Finetuned, Video Backbone=R-152+S3D-HM2021.08 | 68.8 | 27.3 | 56.5 | — | — | 4 | — | |
| UniVL-Jointtraining_loss=LJoint2023.03 | 66.2 | 22.2 | 52.2 | 5 | — | — | — | |
| UniVLapproach=FT-Joint2020.02 | 66.2 | 22.2 | 52.2 | 5 | — | — | — | |
| UniVLPre-training Paradigm=Multi-task Pre-training, Fine-tuning Strategy=FT-Joint2021.05 | 66.2 | 22.2 | 52.2 | — | — | 5 | — | |
| OmniVecModality=Video+Text, Pre-trained=true2025.07 | 64.2 | — | — | — | — | — | — | |
| VideoCLIPzero-shot=true2022.12 | 63.1 | 22.7 | 50.4 | — | — | — | — | |
| TACTAdaptation Dataset=Charades, Zero-shot evaluation=true2023.01 | 62.4 | 22.4 | 49.1 | — | — | — | — | |
| VALUEModality=Vision-only2023.05 | 62.2 | 31.3 | 53 | — | — | — | — | |
| TACTAdaptation Dataset=Charades-Ego, Zero-shot evaluation=true2023.01 | 61.9 | 21.9 | 48.2 | — | — | — | — | |
| BaselineAdaptation Dataset=TEMPO, Zero-shot evaluation=true2023.01 | 61.8 | 21.5 | 48.2 | — | — | — | — | |
| BaselineAdaptation Dataset=Charades, Zero-shot evaluation=true2023.01 | 61.7 | 21.5 | 48.6 | — | — | — | — | |
| BaselineAdaptation Dataset=Charades-Ego, Zero-shot evaluation=true2023.01 | 60.8 | 19.4 | 47.1 | — | — | — | — | |
| VideoCLIPAdaptation Dataset=N/A, Zero-shot evaluation=true2023.01 | 59.9 | 18.2 | 45.5 | — | — | — | — | |
| TACTAdaptation Dataset=TEMPO, Zero-shot evaluation=true2023.01 | 58.7 | 20.4 | 45.1 | — | — | — | — | |
| TACoEvaluation Protocol=Zero-shot, Video Backbone=S3D-HM2021.08 | 55.7 | 19.9 | 43.2 | — | — | 8 | — | |
| TACozero-shot=true2022.12 | 55.7 | 19.9 | 43.2 | — | — | — | — | |
| VideoCoCa (w.o. videos)zero-shot=true2022.12 | 55.2 | 21.7 | 43.9 | — | — | — | — | |
| VideoCoCazero-shot=true2022.12 | 53.3 | 20.3 | 43 | — | — | — | — | |
| TACoLang.=BERT, Video=S3D-HM2021.08 | 53.1 | 16.6 | 40.3 | — | — | 9 | — | |
| CootPre-training Paradigm=Task-specific Alignment Pre-training2021.05 | 52.3 | 16.7 | 40.2 | — | — | 9 | — | |
| BaselineAdaptation Dataset=ActivityNet, Zero-shot evaluation=true2023.01 | 51.4 | 15.6 | 38.8 | — | — | — | — | |
| MIL-NCE2023.03 | 51.2 | 15.1 | 38 | 10 | — | — | — | |
| MIL-NCE2020.02 | 51.2 | 15.1 | 38 | 10 | — | — | — | |
| MIL-NCEEvaluation Protocol=Zero-shot, Video Backbone=S3D-HM2021.08 | 51.2 | 15.1 | 38 | — | — | 10 | — | |
| MIL-NCEModality=VT, Architecture=S3D-G, Training Dataset=HT, Compute=64 TPU, Training Time=3 days, Trainable Backbone=Y2021.04 | 51.2 | 15.1 | 38 | — | — | — | — | |
| MIL-NCEMod=VT, Model=S3D-G, TR=Y, Zero-shot=true2021.04 | 51.2 | 15.1 | 38 | — | — | — | — | |
| TACTAdaptation Dataset=ActivityNet, Zero-shot evaluation=true2023.01 | 49.8 | 16 | 36.9 | — | — | — | — | |
| CoCazero-shot=true2022.12 | 46.4 | 16.8 | 35.8 | — | — | — | — | |
| MMV FACModality=VAT, Architecture=TSM-50x2, Training Dataset=HT+AS, Compute=32 TPU, Training Time=3 days, Trainable Backbone=Y2021.04 | 45.4 | 11.7 | 33.4 | — | — | — | — | |
| MMV FACMod=VAT, Model=TSM-50x2, TR=Y, Zero-shot=true2021.04 | 45.4 | 11.7 | 33.4 | — | — | — | — | |
| MCN (MMS + Cluster + Reconstruct)Loss=MMS + Cluster + Reconstruct2021.04 | 45.2 | — | — | — | — | — | — | |
| MCNModality=VAT, Architecture=R152+RX101, Training Dataset=HT+ImNet+K400, Compute=4 V100, Training Time=2 days, Trainable Backbone=N2021.04 | 45.2 | 18.1 | 35.5 | — | — | — | — | |
| MCN (ours)Mod=VAT, Model=R152+RX101, TR=N, Zero-shot=true2021.04 | 45.2 | 18.1 | 35.5 | — | — | — | — | |
| MMS + ClusterLoss=MMS + Cluster2021.04 | 44.3 | — | — | — | — | — | — | |
| VideoAsMT2020.02 | 43.9 | 11.6 | — | — | — | — | — | |
| VideoAsMTPre-training Paradigm=Pairwise Matching2021.05 | 43.9 | 11.6 | — | — | — | — | — | |
| MMSLoss=MMS2021.04 | 43.7 | — | — | — | — | — | — | |
| MIL-NCEModality=VT, Architecture=I3D-G, Training Dataset=HT, Compute=64 TPU, Training Time=3 days, Trainable Backbone=Y2021.04 | 42 | 11.4 | 30.6 | — | — | — | — | |
| MIL-NCEMod=VT, Model=I3D-G, TR=Y, Zero-shot=true2021.04 | 42 | 11.4 | 30.6 | — | — | — | — | |
| MIL-NCELoss=MIL-NCE2021.04 | 40 | — | — | — | — | — | — | |
| NCELoss=NCE2021.04 | 39.2 | — | — | — | — | — | — | |
| ActBERTFine-tuning=false2020.11 | 38 | 9.6 | 26.7 | 19 | — | — | — | |
| ActBERT2023.03 | 38 | 9.6 | 26.7 | 19 | — | — | — | |
| ActBERT2020.02 | 38 | 9.6 | 26.7 | 19 | — | — | — | |
| ActBERTEvaluation Protocol=Zero-shot, Video Backbone=O-101+ R(2+1)D2021.08 | 38 | 9.6 | 26.7 | — | — | 19 | — | |
| ActBERTPre-training Paradigm=Pairwise Matching2021.05 | 38 | 9.6 | 26.7 | — | — | 19 | — | |
| ActBERTModality=VT, Architecture=R101+Res3D, Training Dataset=HT+VG+K400, Trainable Backbone=N2021.04 | 38 | 9.6 | 26.7 | — | — | — | — | |
| ActBERTMod=VT, Model=R101+Res3D, TR=N, Zero-shot=true2021.04 | 38 | 9.6 | 26.7 | — | — | — | — | |
| TVJEFine-tuning=true2020.11 | 35.3 | 8.2 | 24.5 | 24 | — | — | — | |
| HowTo100MTrainset=PT: HowTo100M FT: YouCook22019.06 | 35.3 | 8.2 | 24.5 | 24 | — | — | — | |
| HowTo100M2020.02 | 35.3 | 8.2 | 24.5 | 24 | — | — | — | |
| TJVEEvaluation Protocol=Finetuned, Video Backbone=R-152+I3D-X1012021.08 | 35.3 | 8.2 | 24.5 | — | — | 24 | — | |
| UniVL(v3)Lang.=BERT, Video=S3D-HM2021.08 | 34.7 | 7.7 | 23.9 | — | — | 21 | — | |
| MIL-NCE*Modality=VT, Architecture=R152+RX101, Training Dataset=HT+ImNet+K400, Compute=4 V100, Training Time=2 days, Trainable Backbone=N2021.04 | 32.3 | 8.1 | 23.3 | — | — | — | — | |
| MIL-NCE*Mod=VT, Model=R152+RX101, TR=N, Zero-shot=true2021.04 | 32.3 | 8.1 | 23.3 | — | — | — | — | |
| VATT + BothPre-training Dataset=HT100M + AudioSet + YT8M, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 31.86 | — | — | — | 29 | — | — | |
| VATT + GRPre-training Dataset=HT100M + AudioSet, Cross-Modality Gradient Realignment (GR)=true2022.11 | 31.65 | — | — | — | 29 | — | — | |
| VATT + CLPre-training Dataset=HT100M + AudioSet + YT8M, Gradient-based Curriculum Learning (CL)=true2022.11 | 31.34 | — | — | — | 31 | — | — | |
| VATT + RW (VT)Pre-training Dataset=HT100M + AudioSet + YT8M, Gradient Re-weighting (RW)=Video-Text2022.11 | 31.07 | — | — | — | 27 | — | — | |
| VATT + BothPre-training Dataset=HT100M + AudioSet, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 30.26 | — | — | — | 31.5 | — | — | |
| VATTPre-training Dataset=HT100M + AudioSet + YT8M2022.11 | 29.66 | — | — | — | 29 | — | — | |
| VATT + GRPre-training Dataset=HT100M + AudioSet + YT8M, Cross-Modality Gradient Realignment (GR)=true2022.11 | 29.56 | — | — | — | 32 | — | — | |
| VATT + CLPre-training Dataset=HT100M + AudioSet, Gradient-based Curriculum Learning (CL)=true2022.11 | 29.17 | — | — | — | 33 | — | — | |
| VATTPre-training Dataset=HT100M + AudioSet2022.11 | 29 | — | — | — | 34 | — | — | |
| VATT + CLPre-training Dataset=HT100M, Gradient-based Curriculum Learning (CL)=true2022.11 | 26.01 | — | — | — | 40 | — | — | |
| VATT + BothPre-training Dataset=HT100M, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 25.87 | — | — | — | 42 | — | — | |
| COOT2023.03 | 25.3 | 16.7 | 40.2 | 9 | — | — | — | |
| VATT + GRPre-training Dataset=HT100M, Cross-Modality Gradient Realignment (GR)=true2022.11 | 24.94 | — | — | — | 47 | — | — | |
| HowTo100MTrainset=HowTo100M2019.06 | 24.8 | 6.1 | 17.3 | 46 | — | — | — | |
| TJVEEvaluation Protocol=Zero-shot, Video Backbone=R-152+I-1012021.08 | 24.8 | 6.1 | 17.3 | — | — | 46 | — | |
| MiechModality=VT, Architecture=R152+RX101, Training Dataset=HT+ImNet+K400, Compute=1 V100, Training Time=1 day, Trainable Backbone=N2021.04 | 24.8 | 6.1 | 17.3 | — | — | — | — | |
| MiechMod=VT, Model=R152+RX101, TR=N, Zero-shot=true2021.04 | 24.8 | 6.1 | 17.3 | — | — | — | — | |
| HowTo100M2023.03 | 24.5 | 8.2 | 35.3 | 24 | — | — | — | |
| FitCLIPzero-shot=true2022.12 | 22.1 | 5.8 | 15.5 | — | — | — | — | |
| TACoLang.=BERT, Video=R-152+I3D-X1012021.08 | 21.7 | 4.9 | 14.7 | — | — | 63 | — | |
| HGLMM2020.11 | 21.6 | 4.6 | 14.3 | 75 | — | — | — | |
| HGLMM FV CCATrainset=YouCook22019.06 | 21.6 | 4.6 | 14.3 | 75 | — | — | — | |
| HGLMM2020.02 | 21.6 | 4.6 | 14.3 | 75 | — | — | — | |
| TVJEFine-tuning=false2020.11 | 21.5 | 4.2 | 13.7 | 65 | — | — | — | |
| HowTo100MTrainset=YouCook22019.06 | 21.5 | 4.2 | 13.7 | 65 | — | — | — |