Video Question Answering on ActivityNet-QA LLaVA-Hound in-domain (test)
68.5AccuracyMM1.5-Video
Evaluation Results
| Method | Links | |
|---|---|---|
| MM1.5-VideoModel Scale=7B, Protocol=SFT2024.09 | 68.5 | |
| MM1.5-VideoModel Scale=3B, Protocol=SFT2024.09 | 67.8 | |
| MM1.5-VideoModel Scale=1B, Protocol=SFT2024.09 | 65.7 | |
| LLAVA-HOUND-SFTModel Scale=7B2024.09 | 62.8 | |
| MM1.5-VideoModel Scale=7B, Protocol=Training-free2024.09 | 52.8 | |
| MM1.5-VideoModel Scale=3B, Protocol=Training-free2024.09 | 51.5 | |
| MM1.5-VideoModel Scale=1B, Protocol=Training-free2024.09 | 49 | |
| Video-LLaVAModel Scale=7B2024.09 | 41.4 | |
| Chat-UniViModel Scale=7B2024.09 | 39.4 | |
| LLaMA-VIDModel Scale=7B2024.09 | 36.5 | |
| Video-ChatGPTModel Scale=7B2024.09 | 34.2 |