Video Captioning on Youcook2
40.4METEORVALUE
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VALUEModality=V+T2021.11 | 40.4 | 12.4 | 18.8 | — | 130.3 | — | |
| MV-GPTPT parts=E+D, Inputs=V+T2022.01 | 27.09 | — | 21.88 | 49.38 | 221 | — | |
| HumanInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 25.9 | — | 15.2 | 45.1 | 3.8 | — | |
| UniVLPT parts=E+D, Inputs=V+T2022.01 | 22.35 | — | 17.35 | 46.52 | 181 | — | |
| UniVLInput=V + T2020.02 | 22.35 | 23.87 | 17.35 | 46.52 | 1.81 | — | |
| MV-GPTPT parts=E+D, Inputs=V2022.01 | 21.43 | — | 16.71 | 41.56 | 153 | — | |
| MV-GPTPT parts=E+D, Inputs=T2022.01 | 20.88 | — | 16.71 | 40.19 | 156 | — | |
| SG-FSCFormerLLM=Vicuna-7B2026.03 | 20.5 | — | — | — | 139.6 | 33.4 | |
| DECEMBERTPT parts=E, Inputs=V+T2022.01 | 20.01 | — | 11.92 | 40.22 | 58 | — | |
| UniVLInput=T2020.02 | 19.39 | 20.32 | 14.7 | 41.1 | 1.51 | — | |
| CootDecoder Type=Extra Decoder2021.05 | 19.34 | 17.62 | 11.09 | 37.63 | — | — | |
| M-MASSPT parts=E+D, Inputs=V+T2022.01 | 18.32 | — | 12.04 | 39.03 | 123 | — | |
| E2vidD6-MASSvid-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 18.32 | — | 12.04 | 39.03 | 1.23 | — | |
| VLMDecoder Type=compact decoder using BERT's LM heads2021.05 | 18.22 | 17.78 | 12.27 | 41.51 | 1.3869 | — | |
| DPCModality=V+T2021.11 | 18.1 | — | 2.8 | — | — | — | |
| HierarQ2025.03 | 18.1 | — | — | — | — | — | |
| DPCInputs=V+T2022.01 | 18.08 | — | 2.76 | — | — | — | |
| DPCInput=V + T2020.02 | 18.08 | 7.6 | 2.76 | — | — | — | |
| DPCInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 18.08 | — | 2.76 | — | — | — | |
| E2vidD2-MASS-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 18.04 | — | 11.38 | 38.67 | 1.19 | — | |
| E2vidD6-MASS-BiD (S3D)Input=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=S3D2020.11 | 18.04 | — | 11.64 | 38.75 | 1.24 | — | |
| E2vidD6-MASS-UniDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 18 | — | 11.39 | 38.71 | 1.22 | — | |
| E2vidD2-MASSdrop-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.99 | — | 11.21 | 38.72 | 1.23 | — | |
| E2vid,D2-MASS-BiDaltInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.85 | — | 11.49 | 38.6 | 1.18 | — | |
| AT+VideoInputs=V+T2022.01 | 17.77 | — | 9.01 | 36.65 | 112 | — | |
| AT+VideoInput=V + T2020.02 | 17.77 | — | 9.01 | 36.65 | 1.12 | — | |
| AT+VideoInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 17.77 | — | 9.01 | 36.65 | 1.12 | — | |
| E2D6-MASS-UniDInput=ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.74 | — | 10.72 | 37.85 | 1.17 | — | |
| E2vidD6-MASSdrop-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.74 | — | 10.45 | 38.82 | 1.22 | — | |
| E2vidD2-MASS-BiD (S3D)Input=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=S3D2020.11 | 17.71 | — | 11.13 | 38.57 | 1.12 | — | |
| E2vidD2-MASSvid-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.71 | — | 11.17 | 38.32 | 1.17 | — | |
| AT+VideoModality=V+T2021.11 | 17.7 | — | 9 | — | 112 | — | |
| E2vidD6-MASS-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.7 | — | 11.47 | 38.8 | 1.25 | — | |
| E2vid,D6-MASS-BiDaltInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.68 | — | 11.07 | 38.43 | 1.22 | — | |
| E2vidD6-MASSalign-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.62 | — | 11.53 | 39.03 | 1.22 | — | |
| DPCModality=V2021.11 | 17.6 | — | 2.2 | — | — | — | |
| MA-LMM2024.04 | 17.6 | — | — | — | 1.312 | — | |
| MA-LMM2025.03 | 17.6 | — | — | — | — | — | |
| MA-LMMLLM=Vicuna-7B2026.03 | 17.6 | — | — | — | 131.2 | 31.5 | |
| UniVLInput=V2020.02 | 17.57 | 16.46 | 11.17 | 40.09 | 1.27 | — | |
| E2vidD2-MASSalign-BiDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.57 | — | 11.54 | 37.7 | 1.15 | — | |
| UniVLDecoder Type=w/ Pre-trained Decoder2021.05 | 17.57 | 16.46 | 11.17 | 40.09 | 1.27 | — | |
| UniVLVideo-level features=Pretrained backbones2022.09 | 17.57 | 16.46 | 11.17 | 40.09 | 1.27 | — | |
| E2D2-MASS-BiDInput=ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.44 | — | 10.84 | 37.2 | 1.13 | — | |
| E2D6-MASS-BiDInput=ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.42 | — | 10.6 | 38.08 | 1.2 | — | |
| E2vidD2-MASS-UniDInput=Video + ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.39 | — | 10.84 | 38.24 | 1.16 | — | |
| GIT2024.04 | 17.3 | — | — | — | 1.298 | — | |
| GIT2025.03 | 17.3 | — | — | — | — | — | |
| E2D2-MASS-UniDInput=ASR, Pretraining=YT8M-cook + Recipe1M, Video features=Compact 2D2020.11 | 17.14 | — | 10.52 | 37.39 | 1.14 | — | |
| ATInputs=T2022.01 | 16.93 | — | 8.55 | 35.54 | 106 | — | |
| ATInput=T2020.02 | 16.93 | — | 8.55 | 35.54 | 1.06 | — | |
| ATInput=ASR, Pretraining=-, Video features=Compact 2D2020.11 | 16.93 | — | 8.55 | 35.54 | 1.06 | — | |
| UniViLM #5Input=Video + ASR, Pretraining=HowTo100M, Video features=Compact 2D2020.11 | 16.93 | — | 10.42 | 38.02 | 1.2 | — | |
| ATVideo-level features=Pretrained backbones2022.09 | 16.93 | — | 8.55 | 35.54 | 1.06 | — | |
| ATModality=T2021.11 | 16.9 | — | 8.5 | — | 106 | — | |
| Video-LLaMA2024.04 | 16.5 | — | — | — | 1.237 | — | |
| E2vidD6-BiD (S3D)Input=Video + ASR, Pretraining=-, Video features=S3D2020.11 | 16.28 | — | 7.91 | 35.23 | 0.93 | — | |
| E2vidD6-BiDInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 16.19 | — | 8.01 | 34.66 | 0.91 | — | |
| E2vidD2-BiD (S3D)Input=Video + ASR, Pretraining=-, Video features=S3D2020.11 | 16.17 | — | 8.04 | 36.01 | 0.96 | — | |
| E2vid,D6-BiDaltInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 16.11 | — | 7.7 | 34.78 | 0.91 | — | |
| MARTInputs=V2022.01 | 15.9 | — | 8 | — | 36 | — | |
| BLIP + HowTo100MVision Encoder=ViT-B, Image-Text Data=6 datasets (CC3M+COCO+VG+SBU+CC12M+LAION), Video-Text Data=HowTo100M, Fine-tuned=true2023.10 | 15.9 | — | 8.6 | 37.1 | 112.9 | — | |
| HowToCaptionVision Encoder=ViT-B, Image-Text Data=6 datasets (CC3M+COCO+VG+SBU+CC12M+LAION), Video-Text Data=HowToCaption, Fine-tuned=true2023.10 | 15.9 | — | 8.8 | 37.3 | 116.4 | — | |
| HowToCaption2025.03 | 15.9 | — | — | — | — | — | |
| MARTInput=Video, Pretraining=-, Video features=Compact 2D2020.11 | 15.9 | — | 8 | — | 0.36 | — | |
| E2vidD2-BiDaltInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.83 | — | 8.12 | 34.83 | 0.93 | — | |
| E2D6-BiDInput=ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.7 | — | 7.9 | 34.86 | 0.93 | — | |
| E2D2-BiDInput=ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.64 | — | 6.85 | 34.26 | 0.91 | — | |
| SWINBERTModality=V2021.11 | 15.6 | 13.8 | 9 | 37.3 | 109 | — | |
| SwinBERT2024.04 | 15.6 | — | — | — | 1.09 | — | |
| SwinBERTVision Encoder=VidSwin-B2023.10 | 15.6 | — | 9 | 37.3 | 109 | — | |
| GIT-2Vision Encoder=DaViT-4.8B, Image-Text Data=12.9B pairs2023.10 | 15.6 | — | 9.4 | 37.5 | 131.2 | — | |
| SwinBERT2025.03 | 15.6 | — | — | — | — | — | |
| E2vidD6-UniDInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.57 | — | 7.61 | 34.28 | 0.89 | — | |
| VideoGLaMMLLM=Phi3-Mini-3.8B2026.03 | 15.4 | — | — | — | 124.3 | 28.7 | |
| UniViLM #2Input=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.38 | — | 8.67 | 35.02 | 1 | — | |
| E2vidD2-BiDInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.36 | — | 8.39 | 34.54 | 0.91 | — | |
| E2D6-UniDInput=ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.29 | — | 7.88 | 34.1 | 0.87 | — | |
| E2D2-UniDInput=ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.15 | — | 7.42 | 33.26 | 0.85 | — | |
| E2vidD2-UniDInput=Video + ASR, Pretraining=-, Video features=Compact 2D2020.11 | 15.11 | — | 7.47 | 34.77 | 0.9 | — | |
| BLIPVision Encoder=ViT-B, Image-Text Data=6 datasets (CC3M+COCO+VG+SBU+CC12M+LAION)2023.10 | 15 | — | 7.9 | 36 | 104.8 | — | |
| OmniVL2022.09 | 14.83 | 12.87 | 8.72 | 36.09 | 1.16 | — | |
| VideoAsMTPT parts=E+D, Inputs=V2022.01 | 13.4 | — | 5.3 | — | — | — | |
| VideoAsMTInput=V2020.02 | 13.4 | — | 5.3 | — | — | — | |
| VideoAsMTDecoder Type=w/ Pre-trained Decoder2021.05 | 13.4 | — | 5.3 | — | — | — | |
| ActBERTPT parts=E, Inputs=V2022.01 | 13.3 | — | 5.41 | 30.56 | 65 | — | |
| ActBERTModality=V2021.11 | 13.3 | 8.6 | 5.4 | — | 65 | — | |
| ActBERTInput=V2020.02 | 13.3 | 8.66 | 5.41 | 30.56 | 0.65 | — | |
| ActBERTDecoder Type=Extra Decoder2021.05 | 13.3 | 8.66 | 5.41 | 30.56 | 0.65 | — | |
| ActBERTVideo-level features=Pretrained backbones2022.09 | 13.3 | 8.66 | 5.41 | 30.56 | 0.65 | — | |
| CBTInput=V2020.02 | 12.97 | — | 5.12 | 30.44 | 0.64 | — | |
| CBTInput=Video, Pretraining=Kinetics + HowTo100M, Video features=Compact 2D2020.11 | 12.97 | — | 5.12 | 30.44 | 0.64 | — | |
| CBTDecoder Type=Extra Decoder2021.05 | 12.97 | — | 5.12 | 30.44 | 0.64 | — | |
| UniViLM #1Input=Video, Pretraining=-, Video features=Compact 2D2020.11 | 12.47 | — | 6.06 | 31.48 | 0.64 | — | |
| GIT-BVision Encoder=ViT-B, Image-Text Data=4 datasets (CC3M+COCO+VG+SBU)2023.10 | 12.2 | — | 5.8 | 31.5 | 80.3 | — | |
| VideoBERTModality=V2021.11 | 11.9 | 7.5 | 4.3 | — | 55 | — | |
| EMTInput=V2020.02 | 11.55 | — | 4.38 | 27.44 | 0.38 | — | |
| EMTInput=Video, Pretraining=-, Video features=Compact 2D2020.11 | 11.55 | — | 4.38 | 27.44 | 0.38 | — | |
| EMTVideo-level features=Pretrained backbones2022.09 | 11.55 | — | 4.38 | 27.44 | 0.38 | — | |
| VideoBERTPT parts=E, Inputs=V2022.01 | 11.01 | — | 4.04 | 27.5 | 49 | — |