Audio Question Answering on ClothoAQA (test)
71.02AccuracyVideoLLaMA2.1-AV
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoLLaMA2.1-AVSize=7B, Training Hours=5k2024.06 | 71.02 | |
| VideoLLaMA2-AVSize=7B, Training Hours=5k2024.06 | 70.11 | |
| Qwen2.5-OmniModel Size=7-8B, Zero-shot evaluation=true2025.12 | 68 | |
| JavisGPTModel Size=7-8B, #Samples=1.5M, Zero-shot evaluation=true2025.12 | 67.3 | |
| VideoLLaMA2.1Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 66.3 | |
| VideoLLaMA2Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 65.1 | |
| Qwen2-AudioModel Size=7-8B, Zero-shot evaluation=true2025.12 | 60.9 | |
| Qwen-AudioSize=7B, Training Hours=137k2024.06 | 57.9 | |
| Qwen-AudioModel Size=7-8B, Zero-shot evaluation=true2025.12 | 57.9 | |
| UnifiedIO-2Model Size=7-8B, #Samples=9.2B, Zero-shot evaluation=true2025.12 | 31.4 | |
| NExT-GPTModel Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 30.9 | |
| InternVideo2_S2Backbone=InternVideo2_S2, Finetuning setting=true2024.03 | 30.14 | |
| MWAFMBackbone=MWAFM, Finetuning setting=true2024.03 | 22.24 | |
| AquaNetBackbone=AquaNet, Finetuning setting=true2024.03 | 14.78 |