Event Captioning on ViTT (test)
51.29CIDErHiCM²
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| HiCM²Pretraining=Yes, Backbone=CLIP2024.12 | 51.29 | 9.66 | 15.07 | 0.86 | |
| Vid2SeqPretraining=Yes, Backbone=CLIP, Reproduced=true2024.12 | 48.84 | 9.51 | 14.99 | 0.71 | |
| Streaming V2SPretraining=Yes, Backbone=CLIP2024.12 | 25.2 | 5.8 | 10 | — |