Video Question Answering on IntentQA
88.6Accuracy (All)Magma-8B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Magma-8BBackbone=LLaMA3-8B, Zero-shot=true2025.02 | 88.6 | — | — | — | — | |
| ProVCALLM/MLLM=GPT-4o2026.04 | 77.7 | — | — | — | — | |
| VideoINSTAArchitecture Type=Bottom-up, zero-shot=true2025.01 | 72.8 | — | — | — | — | |
| VideoINSTAParadigm=Bottom-up, Zero-shot=true2025.01 | 72.8 | — | — | — | — | |
| LVNetArchitecture Type=Bottom-up, zero-shot=true2025.01 | 71.7 | 75 | 74.4 | 62.1 | — | |
| LVNetParadigm=Bottom-up, Zero-shot=true2025.01 | 71.7 | 75 | 74.4 | 62.1 | — | |
| LVNetVideo-level training free=true, Number of captions=122024.06 | 71.7 | — | — | — | — | |
| LVNetLLM/MLLM=GPT-4o2026.04 | 71.7 | — | — | — | — | |
| ENTERArchitecture Type=Top-Down/Modular, zero-shot=true2025.01 | 71.5 | 73.2 | 79.1 | 61.4 | — | |
| ENTERParadigm=Top-Down/Modular, Zero-shot=true2025.01 | 71.5 | 73.2 | 79.1 | 61.4 | — | |
| VideoMindPalace2025.01 | 70.1 | — | — | — | — | |
| Gemini 1.5 FlashArchitecture Type=End-to-End, zero-shot=true2025.01 | 68.8 | 70.8 | 76.2 | 59.2 | — | |
| Gemini 1.5 FlashParadigm=End-to-End, Zero-shot=true2025.01 | 68.8 | 70.8 | 76.2 | 59.2 | — | |
| LLoViArchitecture Type=Bottom-up, zero-shot=true2025.01 | 67.1 | — | — | — | — | |
| LLoViParadigm=Bottom-up, Zero-shot=true2025.01 | 67.1 | — | — | — | — | |
| VideoTreeArchitecture Type=Bottom-up, zero-shot=true2025.01 | 66.9 | — | — | — | — | |
| VideoTreeParadigm=Bottom-up, Zero-shot=true2025.01 | 66.9 | — | — | — | — | |
| VideoTreeVideo-level training free=true, Number of captions=(56)2024.06 | 66.9 | — | — | — | — | |
| VideoTree2025.01 | 66.9 | — | — | — | — | |
| VideoTreeLLM/MLLM=GPT-42026.04 | 66.9 | — | — | — | — | |
| SF-LLAVA-34B2025.01 | 66.5 | — | — | — | — | |
| IG-VLMVideo-level training free=true2024.06 | 65.3 | — | — | — | — | |
| IG-VLMLLM/MLLM=GPT-4V2026.04 | 65.3 | — | — | — | — | |
| TraveLERArchitecture Type=Top-Down/Modular, zero-shot=true2025.01 | 65.2 | 69.9 | 64.7 | 54.4 | — | |
| IG-VLM2025.01 | 64.2 | — | — | — | — | |
| LLoViVideo-level training free=true, Number of captions=902024.06 | 64 | — | — | — | — | |
| LLoVi2025.01 | 64 | — | — | — | — | |
| LLoViLLM/MLLM=GPT-42026.04 | 64 | — | — | — | — | |
| SeViLAArchitecture Type=End-to-End, zero-shot=true2025.01 | 60.9 | — | — | — | — | |
| SeViLAParadigm=End-to-End, Zero-shot=true2025.01 | 60.9 | — | — | — | — | |
| IG-VLMBackbone=Vicuna-7B, Zero-shot=true2025.02 | 60.3 | — | — | — | — | |
| SF-LLaVABackbone=Vicuna-7B, Zero-shot=true2025.02 | 60.1 | — | — | — | — | |
| LangRepoVideo-level training free=true, Number of captions=902024.06 | 59.1 | — | — | — | — | |
| LangRepoLLM/MLLM=Mixtral-8×7B2026.04 | 59.1 | — | — | — | — | |
| VideoSearch-R1Size=2B2026.07 | 57.3 | — | — | — | — | |
| ViperGPTArchitecture Type=Top-Down/Modular, zero-shot=true2025.01 | 54.7 | — | — | — | — | |
| Qwen3-VLSize=2B2026.07 | 39.4 | — | — | — | — | |
| BoxTuningSize=7B2026.04 | — | — | — | — | 76.5 | |
| ENTERZero-shot=true2026.03 | — | — | — | — | 71.5 | |
| IG-VLMZero-shot=true2026.03 | — | — | — | — | 65.3 | |
| LLoViZero-shot=true2026.03 | — | — | — | — | 67.1 | |
| LVNetZero-shot=true2026.03 | — | — | — | — | 71.7 | |
| ObjectMLLMSize=7B2026.04 | — | — | — | — | 75.5 | |
| SeViLAZero-shot=true2026.03 | — | — | — | — | 60.9 | |
| TS-LLAVAZero-shot=true2026.03 | — | — | — | — | 67.9 | |
| VideoAgent2Zero-shot=true2026.03 | — | — | — | — | 73.9 | |
| VideoHV-AgentZero-shot=true2026.03 | — | — | — | — | 75.6 | |
| VideoINSTAZero-shot=true2026.03 | — | — | — | — | 72.8 | |
| VideoLLaMA2Size=7B, Zero-shot=true2026.04 | — | — | — | — | 73.8 | |
| VideoTreeZero-shot=true2026.03 | — | — | — | — | 66.9 |