Audio Visual Question Answering on AVQA (test)
94.3Total AccuracyUniMVU
Evaluation Results
| Method | Links | |
|---|---|---|
| UniMVUSize=7B, Training regime=task-specific training2026.05 | 94.3 | |
| JavisGPTModel Size=7-8B, #Samples=1.5M, Zero-shot evaluation=true2025.12 | 93.8 | |
| PAVE*Size=7B, Evaluation protocol=reproduced baseline2026.05 | 93.4 | |
| UniMVUSize=0.5B, Training regime=task-specific training2026.05 | 92.3 | |
| UniMVU†Size=7B, Training regime=unified multi-task training2026.05 | 92.2 | |
| CAT-FTSize=7B2026.05 | 92 | |
| Qwen2.5-OmniModel Size=7-8B, Zero-shot evaluation=true2025.12 | 91.5 | |
| AV-Master2025.10 | 91.4 | |
| AV-MasterSize=0.5B2026.05 | 91.4 | |
| QSTarPrompting=true2026.01 | 91.2 | |
| UniMVU†Size=0.5B, Training regime=unified multi-task training2026.05 | 91.1 | |
| QSTarPrompting=false2026.01 | 90.9 | |
| TSPMPrompting=true2026.01 | 90.8 | |
| TSPM2025.10 | 90.8 | |
| MCD2025.10 | 90.8 | |
| LLaVA-OV-FT (video-only)Size=7B, Input modalities=video-only2026.05 | 90.8 | |
| PSTPPrompting=false2026.01 | 90.2 | |
| PSTP-Net2025.10 | 90.2 | |
| PSTP-NetSize=0.5B2026.05 | 90.2 | |
| SaSR-Net2025.10 | 89.9 | |
| LLaVA-OV-FT* (video-audio concat)Size=0.5B, Input modalities=video-audio concat, Evaluation protocol=reproduced baseline2026.05 | 89.9 | |
| PAVE*Size=0.5B, Evaluation protocol=reproduced baseline2026.05 | 89.6 | |
| HCRNPrompting=false2026.01 | 89 | |
| HCRNEnsemble=HAVF2025.10 | 89 | |
| ACRTransformerEnsemble=HAVF2025.10 | 87.8 | |
| HGAEnsemble=HAVF2025.10 | 87.7 | |
| PSACPrompting=false2026.01 | 87.4 | |
| PSACEnsemble=HAVF2025.10 | 87.4 | |
| LLaVA-OV-FT (video-only)Size=0.5B, Input modalities=video-only2026.05 | 86.4 | |
| HMEEnsemble=HAVF2025.10 | 85 | |
| LADNetEnsemble=HAVF2025.10 | 84.1 | |
| SWUDI-AZero-shot=true, Method Category=Merging Methods2026.06 | 81.26 | |
| VideoLLaMAModel Size=7-8B, Zero-shot evaluation=true2025.12 | 80.9 | |
| TSV MergingZero-shot=true, Method Category=Merging Methods2026.06 | 80.9 | |
| OptMergeZero-shot=true, Method Category=Merging Methods2026.06 | 80.82 | |
| DAMCZero-shot=true, Method Category=Online Composing2026.06 | 80.78 | |
| NaiveMCZero-shot=true, Method Category=Online Composing2026.06 | 80.26 | |
| VideoZero-shot=true, Method Category=Individual Modalities, Modality=Video2026.06 | 79.2 | |
| Macaw-LLMModel Size=7-8B, Zero-shot evaluation=true2025.12 | 78.7 | |
| AV-LLMModel Size=7-8B, Zero-shot evaluation=true2025.12 | 78.7 | |
| Task ArithmeticZero-shot=true, Method Category=Merging Methods2026.06 | 78.62 | |
| Iso-CZero-shot=true, Method Category=Merging Methods2026.06 | 77.51 | |
| WUDI MergingZero-shot=true, Method Category=Merging Methods2026.06 | 76.86 | |
| TIES MergingZero-shot=true, Method Category=Merging Methods2026.06 | 75.84 | |
| VisionZero-shot=true, Method Category=Individual Modalities, Modality=Vision2026.06 | 75.55 | |
| Weight AverageZero-shot=true, Method Category=Merging Methods2026.06 | 69.39 | |
| UnifiedIO-2Model Size=7-8B, #Samples=9.2B, Zero-shot evaluation=true2025.12 | 61.2 | |
| AudioZero-shot=true, Method Category=Individual Modalities, Modality=Audio2026.06 | 47.57 | |
| NExT-GPTModel Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 25.3 |