Audio-Visual Question Answering on Music-AVQA
81.3AccuracyVideoLLaMA2
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoLLaMA2FLOPs=100, Latency=0.43, Memory=22G2026.01 | 81.3 | |
| VideoLLaMA2 w/ FastAVFLOPs=56, Latency=0.32, Memory=19G2026.01 | 81.2 | |
| VALORLPretraining examples=33.5M, Modality=V+A2023.04 | 78.9 | |
| Q-TriM2026.07 | 77.68 | |
| VALORBPretraining examples=6.5M, Modality=V+A2023.04 | 76.6 | |
| MAVEN2026.07 | 76.21 | |
| VALORLPretraining examples=5.5M, Modality=V2023.04 | 74.8 | |
| MUSIC-AVQAModality=V+A2023.04 | 71.5 | |
| VideoLLaMA22026.07 | 66.85 | |
| Qwen2.5-VLZero-shot=true2026.07 | 57.07 | |
| GPT-4oZero-shot=true2026.07 | 56.06 | |
| TSV MergingEvaluation Protocol=Zero-shot, Composition Strategy=TSV Merging2025.05 | 53.78 | |
| MMER2025.05 | 53.54 | |
| NaiveMCEvaluation Protocol=Zero-shot, Composition Strategy=NaiveMC2025.05 | 53.5 | |
| OptMergeEvaluation Protocol=Zero-shot, Composition Strategy=OptMerge2025.05 | 53.17 | |
| DAMCEvaluation Protocol=Zero-shot, Composition Strategy=DAMC2025.05 | 52.8 | |
| Iso-CEvaluation Protocol=Zero-shot, Composition Strategy=Iso-C2025.05 | 52.77 | |
| WUDI MergingEvaluation Protocol=Zero-shot, Composition Strategy=WUDI Merging2025.05 | 52.43 | |
| Task ArithmeticEvaluation Protocol=Zero-shot, Composition Strategy=Task Arithmetic2025.05 | 52.14 | |
| VisionEvaluation Protocol=Zero-shot, Input Modality=Vision2025.05 | 50.77 | |
| Ties MergingEvaluation Protocol=Zero-shot, Composition Strategy=Ties Merging2025.05 | 50.35 | |
| AVicunaModel size=7B, # Pairs=1.1M, Zero-shot=true2024.11 | 49.6 | |
| LongVALE-LLMModel size=7B, # Pairs=0.7M, Zero-shot=true2024.11 | 49.4 | |
| VideoEvaluation Protocol=Zero-shot, Input Modality=Video2025.05 | 49.02 | |
| CAT-7Bzero-shot=true2024.03 | 48.6 | |
| Weight AverageEvaluation Protocol=Zero-shot, Composition Strategy=Weight Average2025.05 | 47.75 | |
| Video MLLMModality=Video2025.05 | 47.72 | |
| ChatBridge-13Bzero-shot=true2024.03 | 47.6 | |
| OneLLMModel size=7B, # Pairs=1007M, Zero-shot=true2024.11 | 47.6 | |
| AV-LLMModel size=13B, # Pairs=1.6M, Zero-shot=true2024.11 | 45.2 | |
| X-InstructBLIPModel size=13B, # Pairs=32M, Zero-shot=true2024.11 | 44.5 | |
| Vision MLLMModality=Vision2025.05 | 44.06 | |
| OneLLM-7Bzero-shot=true2024.03 | 43 | |
| VideoLLaMAModel size=7B, # Pairs=2.8M, Zero-shot=true2024.11 | 36.6 | |
| PandaGPTModel size=13B, # Pairs=128M, Zero-shot=true2024.11 | 33.7 | |
| Macaw-LLMModel size=7B, # Pairs=0.3M, Zero-shot=true2024.11 | 31.8 | |
| Audio MLLMModality=Audio2025.05 | 30.63 | |
| AudioEvaluation Protocol=Zero-shot, Input Modality=Audio2025.05 | 27.93 |