Video Question Answering on EgoSchema (Full)
77.3AccuracyTASKER
Evaluation Results
| Method | Links | |
|---|---|---|
| TASKER(M)LLM=Qwen3-VL2026.06 | 77.3 | |
| VideoTree(M)LLM=Qwen3-VL2026.06 | 76.7 | |
| Human EvalZero-shot=true2024.05 | 75 | |
| HCQAmode=Zero-shot2025.04 | 75 | |
| VideoAgent(M)LLM=Qwen3-VL2026.06 | 74.6 | |
| GPT-4o2026.05 | 72.2 | |
| LinVT-Qwen2-VLmode=Zero-shot2025.04 | 69.5 | |
| ProVCALLM/MLLM=GPT-4o2026.04 | 69.3 | |
| VideoMultiAgentsmode=Zero-shot2025.04 | 68 | |
| LongVUmode=Zero-shot2025.04 | 67.6 | |
| LongVUTraining Status=Trained2024.05 | 67.6 | |
| Qwen3-VL-4B + MuKVFrames/FPS=0.5fps, # Mem. Tok=5.9K2026.05 | 67 | |
| Qwen3-VL-2B + MuKVFrames/FPS=0.5fps, # Mem. Tok=5.9K2026.05 | 66.3 | |
| Qwen3-VL-4BFrames/FPS=7682026.05 | 65.8 | |
| Gemini 1.5 FlashArchitecture Type=End-to-End, zero-shot=true2025.01 | 65.7 | |
| Gemini 1.5 FlashParadigm=End-to-End, Zero-shot=true2025.01 | 65.7 | |
| VideoINSTAArchitecture Type=Bottom-up, zero-shot=true2025.01 | 65 | |
| VideoINSTAParadigm=Bottom-up, Zero-shot=true2025.01 | 65 | |
| Qwen3-VL-2BFrames/FPS=7682026.05 | 65 | |
| Lifelong Memorymode=Zero-shot2025.04 | 64.7 | |
| Qwen2.5-VL-3BFrames/FPS=7682026.05 | 64.4 | |
| VideoLLaMA2mode=Zero-shot2025.04 | 63.9 | |
| VideoLLaMA 2Training Status=Trained2024.05 | 63.9 | |
| AKEYS(M)LLM=GPT-4o2025.03 | 63.6 | |
| TASKER(M)LLM=GPT-4o2026.06 | 63.6 | |
| InternVL-3 + A.I.R.VLM Size=8B, #Frames=≤ 322025.10 | 63.3 | |
| LLaVA-OV-7B + MuKVFrames/FPS=0.5fps, # Mem. Tok=5.9K2026.05 | 63.3 | |
| Gemini-1.5-ProCore LLMs=Gemini-1.5-Pro, Zero-shot=true2024.05 | 63.2 | |
| AKEYS(M)LLM=GPT-42025.03 | 63.1 | |
| TASKER(M)LLM=GPT-42026.06 | 63.1 | |
| LLaVA-OV-7B + LiveVLMFrames/FPS=0.5/0.2 fps2026.05 | 63 | |
| LLaVA-OV-7B + StreamMemFrames/FPS=0.5/0.2 fps, # Mem. Tok=6K2026.05 | 63 | |
| Qwen2.5-VL-3B + MuKVFrames/FPS=0.5 fps, # Mem. Tok=5.9K2026.05 | 63 | |
| InternVL-3VLM Size=8B, #Frames=322025.10 | 62.5 | |
| LifelongMemorycategory=LLM-based video agents, Core VLMs=LaViLa, Core LLMs=GPT-4, in-domain training (Ego4D)=true, Zero-shot=true2024.05 | 62.4 | |
| Qwen2.5-VL-3B + StreamMemFrames/FPS=4.0/0.5 fps, # Mem. Tok=6K2026.05 | 62.2 | |
| LifelongMemoryEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 62.1 | |
| LifelongMemoryVideo-level training free=true, Number of captions=902024.06 | 62.1 | |
| LLaVA-OV-72BTraining Status=Trained2024.05 | 62 | |
| Qwen2.5-VL-3B + InfiniPot-VFrames/FPS=768, # Mem. Tok=6K2026.05 | 61.8 | |
| Tarsier-34BMode=Zero-shot2024.06 | 61.7 | |
| TarsierEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=34B2024.03 | 61.7 | |
| TarsierVideo-level training free=false2024.06 | 61.7 | |
| Tarsiermode=Zero-shot2025.04 | 61.7 | |
| Tarsier-34BTraining Status=Trained2024.05 | 61.7 | |
| InternVL-3 + MDP3VLM Size=8B, #Frames=322025.10 | 61.6 | |
| LLaVA-OneVision + A.I.R.VLM Size=7B, #Frames=≤ 322025.10 | 61.4 | |
| VideoTreeMode=Zero-shot2024.06 | 61.1 | |
| LVNetEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=<1.8T2024.03 | 61.1 | |
| VideoTreeEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 61.1 | |
| VideoTreeArchitecture Type=Bottom-up, zero-shot=true2025.01 | 61.1 | |
| LVNetArchitecture Type=Bottom-up, zero-shot=true2025.01 | 61.1 | |
| VideoTreeParadigm=Bottom-up, Zero-shot=true2025.01 | 61.1 | |
| LVNetParadigm=Bottom-up, Zero-shot=true2025.01 | 61.1 | |
| LVNet(M)LLM=GPT-4o2025.03 | 61.1 | |
| VideoTree(M)LLM=GPT-42025.03 | 61.1 | |
| VideoTreeVideo-level training free=true, Number of captions=62.42024.06 | 61.1 | |
| LVNetVideo-level training free=true, Number of captions=122024.06 | 61.1 | |
| VideoTreemode=Zero-shot2025.04 | 61.1 | |
| LVNetmode=Zero-shot2025.04 | 61.1 | |
| VIDEOTREETraining Status=Training-free2024.05 | 61.1 | |
| VideoTreeVLM Size=GPT42025.10 | 61.1 | |
| VideoTreeLLM/MLLM=GPT-42026.04 | 61.1 | |
| LVNetLLM/MLLM=GPT-4o2026.04 | 61.1 | |
| LVNet(M)LLM=GPT-4o2026.06 | 61.1 | |
| VideoTree(M)LLM=GPT-42026.06 | 61.1 | |
| DrVideoVLM Size=GPT42025.10 | 61 | |
| DrVideoLLM/MLLM=GPT-42026.04 | 61 | |
| LLaVA-OneVision + BOLTVLM Size=7B, #Frames=322025.10 | 60.7 | |
| LLaVA-OV-7B + ReKVFrames/FPS=0.5 fps, # Mem. Tok=5.9K2026.05 | 60.7 | |
| LLaVA-OneVision 32 frames + BOLTLLM Size=7B, Training Free=true, Frames=322025.03 | 60.66 | |
| LLaVA-OneVision 32 framesLLM Size=7B, Training Free=false, Frames=322025.03 | 60.36 | |
| LLaVA-OneVision + MDP3VLM Size=7B, #Frames=322025.10 | 60.3 | |
| VideoAgentcategory=LLM-based video agents, Core VLMs=Video-LLaVA [32], Core LLMs=GPT-4, citation=[13], Zero-shot=true2024.05 | 60.2 | |
| InternVideo2Evaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=13B2024.03 | 60.2 | |
| VideoAgent [6](M)LLM=GPT-42025.03 | 60.2 | |
| InternVideo2Video-level training free=false2024.06 | 60.2 | |
| VideoChat2Training Status=Trained2024.05 | 60.2 | |
| LLaVA-OneVisionVLM Size=7B, #Frames=322025.10 | 60.2 | |
| VideoAgentLLM/MLLM=GPT-4, Reference=[28]2026.04 | 60.2 | |
| VideoAgent(M)LLM=GPT-42026.06 | 60.2 | |
| LLaVA-OV-7BFrames/FPS=322026.05 | 60.1 | |
| VideoChat-TLLM Size=7B2024.10 | 60 | |
| LLaVA-OneVision 16 frames + BOLTLLM Size=7B, Training Free=true, Frames=162025.03 | 59.86 | |
| IG-VLMEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 59.8 | |
| IG-VLMVideo-level training free=true2024.06 | 59.8 | |
| LLaVA-OneVision 16 framesLLM Size=7B, Training Free=false, Frames=162025.03 | 59.49 | |
| LLaVA-OneVision 8 frames + BOLTLLM Size=7B, Training Free=true, Frames=82025.03 | 59.23 | |
| LLaVA-OneVision 8 framesLLM Size=7B, Training Free=false, Frames=82025.03 | 59.17 | |
| QwenVL-2.5 + A.I.R.VLM Size=7B, #Frames=≤ 322025.10 | 58.8 | |
| LifelongMemory(M)LLM=GPT-42025.03 | 58.6 | |
| LifelongMemory(M)LLM=GPT-42026.06 | 58.6 | |
| LongVUFrames/FPS=400/1fps2026.05 | 58.2 | |
| QwenVL-2.5VLM Size=7B, #Frames=322025.10 | 57.6 | |
| ProViQEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=175B2024.03 | 57.1 | |
| ProViQVideo-level training free=true, Number of captions=502024.06 | 57.1 | |
| QwenVL-2.5 + MDP3VLM Size=7B, #Frames=322025.10 | 56.8 | |
| MovieChat+LLM Size=7B2024.10 | 56.4 | |
| Gemini 1.0 Pro2024.03 | 55.7 | |
| Gemini 1.0 Pro2024.03 | 55.7 |