Audio Description on MAD-Eval (test)
27.3CIDErDistinctAD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DistinctADTraining Approach=Partial-fine-tuning, VLM=CLIP_AD-B16, LLM=LLaMA3-8B2024.11 | 27.3 | 17.6 | 8.3 | 56 | |
| DistinctADTraining Approach=Partial-fine-tuning, VLM=CLIP_AD-B16, LLM=LLaMA2-7B2024.11 | 27 | 17.2 | 8.2 | 55.6 | |
| DistinctADTraining Approach=Partial-fine-tuning, VLM=CLIP_AD-B32, LLM=GPT-22024.11 | 25.5 | 16.4 | 7.4 | 51.7 | |
| DistinctADTraining Approach=Partial-fine-tuning, VLM=CLIP-B32, LLM=GPT-22024.11 | 24.5 | 15.4 | 6.7 | 49.8 | |
| MovieSeqTraining Approach=Partial-fine-tuning, Pub.=ECCV'24, VLM=CLIP-B16, LLM=LLaMA2-7B, LoRA fine-tuning=true2024.11 | 24.4 | 15.5 | 7 | 51.6 | |
| AutoAD-IIITraining Approach=Partial-fine-tuning, Pub.=CVPR'24, VLM=EVA-CLIP, LLM=LLaMA2-7B2024.11 | 24 | — | — | 52.8 | |
| AutoAD-IIITraining Approach=Partial-fine-tuning, Pub.=CVPR'24, VLM=EVA-CLIP, LLM=OPT-2.7B2024.11 | 22.8 | — | — | 52 | |
| AutoAD-ZeroTraining Approach=Training-free, Pub.=ACCV'24, VLM=VideoLLaMA2-7B, LLM=LLaMA3-8B2024.11 | 22.4 | — | — | — | |
| LLM-ADTraining Approach=Training-free, Pub.=ArXiv'24, VLM=GPT-4V2024.11 | 20.5 | 13.5 | — | — | |
| AutoAD-IITraining Approach=Partial-fine-tuning, Pub.=ICCV'23, VLM=CLIP-B32, LLM=GPT-22024.11 | 19.5 | 13.4 | — | 50.8 | |
| AutoAD-ITraining Approach=Partial-fine-tuning, Pub.=CVPR'23, VLM=CLIP-B32, LLM=GPT-22024.11 | 14.3 | 11.9 | 4.4 | 42.1 | |
| MM-NarratorTraining Approach=Training-free, Pub.=CVPR'24, VLM=CLIP-L14, LLM=GPT-42024.11 | 13.9 | 13.4 | 5.2 | 49 | |
| CapDecTraining Approach=Partial-fine-tuning, Pub.=ArXiv'222024.11 | 6.7 | 8.2 | 1.4 | — | |
| MM-VidTraining Approach=Training-free, Pub.=ArXiv'23, VLM=GPT-4V2024.11 | 6.1 | 9.8 | 3.8 | 46.1 | |
| ClipCapTraining Approach=Partial-fine-tuning, Pub.=ArXiv'21, VLM=CLIP-B32, LLM=GPT-22024.11 | 4.4 | 8.5 | 1.1 | — | |
| VLogTraining Approach=Training-free, LLM=GPT-42024.11 | 1.3 | 7.5 | 2.1 | 42.3 |