Video Captioning on TVC (val)
66.1CIDErGIT2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GIT22022.05 | 66.1 | 16.9 | 37.2 | 19.4 | — | |
| GIT22022.05 | 66.1 | 16.9 | 37.2 | 19.4 | — | |
| CLIP4Caption++model ensemble=true, with subtitle=true2022.05 | 66 | 15 | 36.9 | — | — | |
| CLIP4Caption++Subtitle as additional input=true, Model ensemble=true2022.05 | 66 | 15 | 36.9 | — | — | |
| GIT2022.05 | 63 | 16.2 | 36.7 | 18.9 | — | |
| GIT2022.05 | 63 | 16.2 | 36.7 | 18.9 | — | |
| GIT-Large2022.05 | 55.7 | 14.9 | 35.4 | 18 | — | |
| GITBackbone=Large2022.05 | 55.7 | 14.9 | 35.4 | 18 | — | |
| SwinBERT2022.05 | 55.4 | 14.5 | 36.1 | 18.5 | — | |
| SwinBERT2022.05 | 55.4 | 14.5 | 36.1 | 18.5 | — | |
| VALUEwith subtitle=true2022.05 | 50.5 | 11.6 | 33.9 | 17.6 | — | |
| VALUESubtitle as additional input=true2022.05 | 50.5 | 11.6 | 33.9 | 17.6 | — | |
| HEROwith subtitle=true2022.05 | 49.9 | 12.3 | 34.1 | 17.6 | — | |
| HEROSubtitle as additional input=true2022.05 | 49.9 | 12.3 | 34.1 | 17.6 | — | |
| GIT-Base2022.05 | 47.3 | 13 | 33.2 | 16.6 | — | |
| GITBackbone=Base2022.05 | 47.3 | 13 | 33.2 | 16.6 | — | |
| MMTwith subtitle=true2022.05 | 45.3 | 10.8 | 32.8 | 16.9 | — | |
| MMTSubtitle as additional input=true2022.05 | 45.3 | 10.8 | 32.8 | 16.9 | — | |
| HERO# Samples for Multimodal Pretext=7.6M2021.01 | 0.505 | 0.123 | 0.341 | 0.175 | — | |
| VX2TEXT# Samples for Multimodal Pretext=02021.01 | 0.482 | 0.116 | 0.328 | 0.172 | — | |
| MMT# Samples for Multimodal Pretext=02021.01 | 0.444 | 0.105 | 0.324 | 0.166 | — | |
| HERO# Samples for Multimodal Pretext=02021.01 | 0.436 | 0.107 | 0.327 | 0.164 | — |