Synchronized Video Storytelling on E-SyncVidStory (test)
33.3CIDErVideoNarrator
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VideoNarratorCategory=Fine-tuned End2end MLLMs2024.05 | 33.3 | 9.1 | 21 | 9.9 | 6.3 | 4.4 | |
| VideoNarratorAblation=w/o Visual Compression2024.05 | 29.6 | 8.7 | 20.4 | 9.2 | 5.7 | 3.8 | |
| VideoNarratorAblation=w/o Video Clip Position2024.05 | 29.5 | 8.7 | 20.1 | 9.1 | 5.6 | 3.9 | |
| VTimeLLMCategory=Fine-tuned End2end MLLMs, finetuned=true2024.05 | 28 | 8.6 | 19.8 | 9.3 | 5.6 | 3.8 | |
| VideoNarratorAblation=w/o Storyline2024.05 | 26.2 | 8.2 | 19.4 | 9.4 | 5.9 | 3.6 | |
| VideoNarratorAblation=w/o Visual Memory2024.05 | 22.3 | 8.2 | 19.2 | 8.1 | 4.7 | 3.1 | |
| LLaVA-1.5+GPT-3.5Category=Multi-Model Pipelines, Protocol=Few-Shot2024.05 | 18.6 | 8.3 | 20 | 8.5 | 4.9 | 3.2 | |
| LLaVA-1.5+GPT-3.5Category=Multi-Model Pipelines, Protocol=Zero-Shot2024.05 | 15 | 8.5 | 18.4 | 7.6 | 4.3 | 2.7 | |
| VideoNarratorAblation=w/o LLM LoRa finetune2024.05 | 10.5 | 6.9 | 16.5 | 5.4 | 2.8 | 1.7 | |
| VTimeLLMCategory=End2end MLLMs2024.05 | 8.4 | 8.6 | 14.9 | 6.5 | 3.9 | 2.6 | |
| Video-LLaVACategory=End2end MLLMs2024.05 | 6.1 | 8.3 | 15.1 | 6 | 3.3 | 1.8 | |
| Video-ChatGPTCategory=End2end MLLMs2024.05 | 4.1 | 8.6 | 13.8 | 5.5 | 2.9 | 1.7 |