Video Captioning on YouCook II (val)
221CIDErMV-GPT
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MV-GPTwith the subtitle as additional input=true2022.05 | 221 | — | 21.9 | 27.1 | — | — | — | 49.4 | — | — | — | — | |
| MV-GPTwith subtitle=true2022.05 | 221 | — | 21.9 | 27.1 | — | — | — | 49.4 | — | — | — | — | |
| UniVLwith the subtitle as additional input=true2022.05 | 181 | — | 17.4 | 22.4 | — | — | — | 46.5 | — | — | — | — | |
| UniVLwith subtitle=true2022.05 | 181 | — | 17.4 | 22.4 | — | — | — | 46.5 | — | — | — | — | |
| SotA2022.04 | 138.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Restricted SotA+ (Miech et al.)Evaluation Protocol=Fine-tuned2022.04 | 138.7 | — | — | — | — | — | — | — | — | — | — | — | |
| VLMEvaluation Protocol=finetuning2022.12 | 138.7 | — | 12.3 | — | — | — | — | 41.5 | — | — | — | — | |
| HierarQLLM Finetuning=true2025.03 | 136.1 | — | — | — | — | — | — | — | — | — | — | — | |
| HierarQLLM Finetuning=false2025.03 | 134.4 | — | — | — | — | — | — | — | — | — | — | — | |
| TextKGInput=V+S, Evaluation Mode=micro-level2023.03 | 133 | — | — | 18.4 | — | — | — | 40.2 | — | — | 11.7 | — | |
| MA-LMM2024.11 | 131.2 | — | — | 17.6 | — | — | — | — | — | — | — | — | |
| MA-LMM2025.03 | 131.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GIT22022.05 | 131.2 | — | 9.4 | 15.6 | — | — | — | 37.5 | — | — | — | — | |
| GIT22022.05 | 131.2 | — | 9.4 | 15.6 | — | — | — | 37.5 | — | — | — | — | |
| GIT2Evaluation Protocol=finetuning2022.12 | 131.2 | — | 9.4 | — | — | — | — | 37.5 | — | — | — | — | |
| VALUEwith the subtitle as additional input=true2022.05 | 130.3 | — | 12.4 | 18.8 | — | — | — | 40.4 | — | — | — | — | |
| VALUEwith subtitle=true2022.05 | 130.3 | — | 12.4 | 18.8 | — | — | — | 40.4 | — | — | — | — | |
| VALUEInput=V+S, Evaluation Mode=micro-level2023.03 | 130 | — | — | 18.8 | — | — | — | 40.4 | — | — | 12.4 | — | |
| GIT2024.11 | 129.8 | — | — | 17.3 | — | — | — | — | — | — | — | — | |
| GIT2025.03 | 129.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GIT2022.05 | 129.8 | — | 10.3 | 17.3 | — | — | — | 39.8 | — | — | — | — | |
| GIT2022.05 | 129.8 | — | 10.3 | 17.3 | — | — | — | 39.8 | — | — | — | — | |
| VideoCoca2024.11 | 128 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoCoCaEvaluation Protocol=finetuning, Transcripts in encoder=No2022.12 | 128 | — | 14.2 | — | — | — | — | 37.7 | — | — | — | — | |
| UniVL2024.11 | 127 | — | — | — | — | — | — | — | — | — | — | — | |
| UniVLEvaluation Protocol=finetuning2022.12 | 127 | — | 11.2 | — | — | — | — | 40.1 | — | — | — | — | |
| AdaCM²2024.11 | 125.6 | — | — | 17.6 | — | — | — | — | — | — | — | — | |
| VideoLLaMA2024.11 | 123.7 | — | — | 16.5 | — | — | — | — | — | — | — | — | |
| Flamingomode=Fine-tuned2022.04 | 118.6 | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Fine-tuned2022.04 | 118.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingo2022.05 | 118.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingo2022.05 | 118.6 | — | — | — | — | — | — | — | — | — | — | — | |
| HowToCaption2025.03 | 116.4 | — | — | — | — | — | — | — | — | — | — | — | |
| OmniVLEvaluation Protocol=finetuning2022.12 | 116 | — | 8.7 | — | — | — | — | 36.1 | — | — | — | — | |
| UnivlInput=V+S, Evaluation Mode=micro-level2023.03 | 115 | — | — | 16.3 | — | — | — | 37.4 | — | — | 9.5 | — | |
| ATInput=V+S, Evaluation Mode=micro-level2023.03 | 112 | — | — | 17.8 | — | — | — | 36.7 | — | — | 9 | — | |
| SwinBERTInput=V, Evaluation Mode=micro-level2023.03 | 109 | — | — | 15.6 | — | — | — | 37.3 | — | — | 9 | — | |
| SwinBERT2024.11 | 109 | — | — | 15.6 | — | — | — | — | — | — | — | — | |
| SwinBERT2025.03 | 109 | — | — | — | — | — | — | — | — | — | — | — | |
| SwinBERT2022.05 | 109 | — | 9 | 15.6 | — | — | — | 37.3 | — | — | — | — | |
| SwinBERT2022.05 | 109 | — | 9 | 15.6 | — | — | — | 37.3 | — | — | — | — | |
| MV-GPTInput=V+S, Evaluation Mode=micro-level2023.03 | 103 | — | — | 17.6 | — | — | — | 35.5 | — | — | 13.3 | — | |
| GITLmodel size=Large2022.05 | 98.3 | — | 7.5 | 14.4 | — | — | — | 34.9 | — | — | — | — | |
| GIT-Large2022.05 | 98.3 | — | 7.5 | 14.4 | — | — | — | 34.9 | — | — | — | — | |
| FlamingoFew-shot=32 shots2022.04 | 86.8 | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoShots=32, Evaluation Protocol=In-context learning2022.04 | 86.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GITBmodel size=Base2022.05 | 80.3 | — | 5.8 | 12.2 | — | — | — | 31.5 | — | — | — | — | |
| GIT-Base2022.05 | 80.3 | — | 5.8 | 12.2 | — | — | — | 31.5 | — | — | — | — | |
| TextKG2023.03 | 75.9 | — | 14 | 22.1 | — | — | — | — | — | 2.8 | — | — | |
| ActBERTInput=V, Evaluation Mode=micro-level2023.03 | 65 | — | — | 14.3 | — | — | — | 30.6 | — | — | 5.4 | — | |
| ActBERT2022.05 | 65 | — | 5.4 | 13.3 | — | — | — | — | — | — | — | — | |
| ActBERT2022.05 | 65 | — | 5.4 | 13.3 | — | — | — | — | — | — | — | — | |
| Flamingo-80BEvaluation Protocol=zero-shot2022.12 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | |
| COOT2023.03 | 57.2 | — | 11.3 | 19.9 | — | — | — | — | — | 6.7 | — | — | |
| Van-Trans+COOT2023.03 | 55.6 | — | 11.1 | 19.8 | — | — | — | — | — | 5.7 | — | — | |
| VideoBERT2022.05 | 55 | — | 4.3 | 11.9 | — | — | — | — | — | — | — | — | |
| VideoBERT2022.05 | 55 | — | 4.3 | 11.9 | — | — | — | — | — | — | — | — | |
| Vid2SeqBackbone=V (CLIP), Proposals=Learnt2023.02 | 50.1 | — | — | 24 | — | — | — | — | — | — | — | — | |
| VideoBERT+S3DInput=V, Evaluation Mode=micro-level2023.03 | 50 | — | — | 11.9 | — | — | — | 28.8 | — | — | 4.3 | — | |
| VLCapInput=C3D + Language2022.06 | 49.41 | — | 9.56 | 17.95 | — | — | 5.16 | 35.17 | 67.97 | — | — | — | |
| VideoBERTInput=V, Evaluation Mode=micro-level2023.03 | 49 | — | — | 11 | — | — | — | 27.5 | — | — | 4 | — | |
| VLTinTInput=C3D/Ling2022.11 | 48.7 | — | 9.4 | 17.94 | — | — | 4.29 | 34.55 | — | — | — | — | |
| MART w/ COOTInput=COOT2022.06 | 46.06 | — | 9.44 | 18.17 | — | — | 6.3 | — | — | — | — | — | |
| MART^COOTVenue=NIPS, Input=COOT2022.11 | 46.06 | — | 9.44 | 18.17 | — | — | 6.3 | — | — | — | — | — | |
| GPaSInput=Res2002022.06 | 41.44 | — | 1.64 | 12.2 | — | — | — | 27.98 | — | — | — | — | |
| Vanilla TransformerInput=Res200 + Flow2022.06 | 38 | — | 4.38 | 11.55 | — | — | — | — | — | — | — | — | |
| Vanilla Trans.Venue=CVPR, Input=Res200/Flow2022.11 | 38 | — | 4.38 | 11.55 | — | — | 4.39 | — | — | — | — | — | |
| Masked TransInput=V, Evaluation Mode=micro-level2023.03 | 38 | — | — | 11.6 | — | — | — | 27.4 | — | — | 3.8 | — | |
| MARTsentence-level recurrence=true2020.05 | 35.74 | — | 8 | 15.9 | — | 4.39 | — | — | — | — | — | — | |
| MARTInput=Res200 + Flow2022.06 | 35.74 | — | 8 | 15.9 | — | — | 4.39 | — | — | — | — | — | |
| MARTVenue=ACL, Input=Res200/Flow2022.11 | 35.74 | — | 8 | 15.9 | — | — | 4.39 | — | — | — | — | — | |
| MARTBackbone=V (ResNet-200) + F, Proposals=Ground Truth2023.02 | 35.7 | — | — | 15.9 | — | — | — | — | — | — | — | — | |
| MART2023.03 | 35.7 | — | 8 | 16 | — | — | — | — | — | 4.4 | — | — | |
| VideoCoCaEvaluation Protocol=zero-shot, Transcripts in encoder=No2022.12 | 34.3 | — | 7.7 | — | — | — | — | 16.5 | — | — | — | — | |
| VTransformerBackbone=V (ResNet-200) + F, Proposals=Ground Truth2023.02 | 32.3 | — | — | 15.7 | — | — | — | — | — | — | — | — | |
| Van-Trans2023.03 | 32.3 | — | 7.6 | 15.7 | — | — | — | — | — | 7.8 | — | — | |
| VTransformersentence-level recurrence=false2020.05 | 32.26 | — | 7.62 | 15.65 | — | 7.83 | — | — | — | — | — | — | |
| S3DInput=V, Evaluation Mode=micro-level2023.03 | 31 | — | — | 9.5 | — | — | — | 26.1 | — | — | 3.2 | — | |
| Katna + LLaVA + LLMSampling Strategy=Katna key frame sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 28.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer-XLBackbone=V (ResNet-200) + F, Proposals=Ground Truth2023.02 | 26.4 | — | — | 14.8 | — | — | — | — | — | — | — | — | |
| Transformer-XLsentence-level recurrence=true2020.05 | 26.35 | — | 6.56 | 14.76 | — | 6.3 | — | — | — | — | — | — | |
| Regular Sampling + LLaVA + LLMSampling Strategy=Regular Sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 26.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Trans-XL2023.03 | 26.1 | — | 6.6 | 11.8 | — | — | — | — | — | 6.3 | — | — | |
| Random Sampling + LLaVA + LLMSampling Strategy=Random Sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 26 | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer-XLRGsentence-level recurrence=true2020.05 | 25.93 | — | 6.63 | 14.74 | — | 6.03 | — | — | — | — | — | — | |
| Trans-XLRG2023.03 | 25.9 | — | 6.6 | 14.7 | — | — | — | — | — | 6 | — | — | |
| Multiclips: Clip Sampling + Video-LLaVA + LLMSampling Strategy=Clip Sampling, LMM=Video-LLaVA, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 25.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaVAArchitecture=Early fusion2026.01 | 19.9 | — | — | — | — | — | — | — | — | — | — | — | |
| ActBERT2020.11 | 0.65 | 8.66 | 5.41 | 13.3 | 30.56 | — | — | — | — | — | — | — | |
| VideoBERT + S3DFusion=Ensemble with S3D, Segments=ground truth video segments2019.04 | 0.55 | 7.59 | 4.33 | 11.94 | 28.8 | — | — | — | — | — | — | — | |
| VideoBert + S3D2020.11 | 0.55 | 7.59 | 4.33 | 11.94 | 28.8 | — | — | — | — | — | — | — | |
| VideoBERTSegments=ground truth video segments2019.04 | 0.49 | 6.8 | 4.04 | 11.01 | 27.5 | — | — | — | — | — | — | — | |
| VideoBert2020.11 | 0.49 | 6.8 | 4.04 | 11.01 | 27.5 | — | — | — | — | — | — | — | |
| VideoBERT (video only)Modality=video only, Segments=ground truth video segments2019.04 | 0.47 | 6.33 | 3.81 | 10.81 | 27.14 | — | — | — | — | — | — | — | |
| Zhou et al.Segments=ground truth video segments2019.04 | 0.38 | 7.53 | 3.84 | 11.55 | 27.44 | — | — | — | — | — | — | — | |
| Zhou et al.2020.11 | 0.38 | 7.53 | 3.84 | 11.55 | 27.44 | — | — | — | — | — | — | — | |
| S3DSegments=ground truth video segments2019.04 | 0.31 | 6.12 | 3.24 | 9.52 | 26.09 | — | — | — | — | — | — | — | |
| S3D2020.11 | 0.31 | 6.12 | 3.24 | 9.52 | 26.09 | — | — | — | — | — | — | — | |
| DPCInput=V+S, Evaluation Mode=micro-level2023.03 | — | — | — | 18.1 | — | — | — | — | — | — | 2.8 | — | |
| Human2020.05 | — | — | — | — | — | 1.27 | — | — | — | — | — | — |