Audio-Visual Question Answering on AVSD
54.8AccuracyLongVALE-LLM
Evaluation Results
| Method | Links | |
|---|---|---|
| LongVALE-LLMModel size=7B, # Pairs=0.7M, Zero-shot=true2024.11 | 54.8 | |
| AVicunaModel size=7B, # Pairs=1.1M, Zero-shot=true2024.11 | 53.1 | |
| AV-LLMModel size=13B, # Pairs=1.6M, Zero-shot=true2024.11 | 52.6 | |
| VideoLLaMAModel size=7B, # Pairs=2.8M, Zero-shot=true2024.11 | 36.7 | |
| Macaw-LLMModel size=7B, # Pairs=0.3M, Zero-shot=true2024.11 | 34.3 | |
| PandaGPTModel size=13B, # Pairs=128M, Zero-shot=true2024.11 | 26.1 |