Long-video Question Answering on MLVU
79.5M-AvgInternVL-3-78B
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL-3-78BModel Type=Open-Source 72B-78B Model2025.06 | 79.5 | |
| GPT-5LLM=N/A, # Frames=N/A2025.12 | 77.3 | |
| GOPAgenType=Agentic video understanding framework2026.06 | 77.3 | |
| VideoChat-A1Backbone=InternVideo2.5-8B2025.06 | 76.2 | |
| VideoLucyType=Agentic video understanding framework2026.06 | 76.1 | |
| InternVL2.5-78B2025.01 | 75.7 | |
| InternVL-2.5-78BModel Type=Open-Source 72B-78B Model2025.06 | 75.7 | |
| InternVL2.5-78BAccess=Open-Source2026.06 | 75.7 | |
| VideoChat-A1Backbone=InternVL2.5-8B2025.06 | 75.1 | |
| Qwen2.5-VL-72BAccess=Open-Source2026.06 | 74.6 | |
| LLaVA-Video-72B2025.01 | 74.4 | |
| VideoRAG-72BModel Type=Agent Based Model2025.06 | 73.8 | |
| InternVideo-2.5-8BModel Type=Open-Source 7B-8B Model2025.06 | 72.8 | |
| VideoRAGType=Agentic video understanding framework2026.06 | 72.4 | |
| LLaVA-Video + OneClip-RAGLLM=7B, # Frames=642025.12 | 72.3 | |
| VideoChat-A1Backbone=Qwen2.5-VL-7B2025.06 | 71.9 | |
| AKSLLM=7B, # Frames=642025.12 | 71.4 | |
| InternVL-3-8BModel Type=Open-Source 7B-8B Model2025.06 | 71.4 | |
| LLaVA-VideoLLM=7B, # Frames=642025.12 | 71.1 | |
| LLaVA-Video-7B2025.01 | 70.8 | |
| LLaVA-Video-7BModel Type=Open-Source 7B-8B Model2025.06 | 70.8 | |
| Qwen2.5-VLLLM=7B, # Frames=322025.12 | 70.2 | |
| InternVL3.5LLM=8B, # Frames=642025.12 | 70.2 | |
| NVILALLM=8B, # Frames=2562025.12 | 70.1 | |
| ByteVideoLLMLLM=14B, # Frames=2562025.12 | 70.1 | |
| InternVL2.5-8B2025.01 | 68.9 | |
| InternVL-2.5-8BModel Type=Open-Source 7B-8B Model2025.06 | 68.9 | |
| LLaVA-OneVision-72BModel Type=Open-Source 72B-78B Model2025.06 | 68 | |
| Tarsier2-7BNumber of frames=256f2025.01 | 67.9 | |
| LongVULLM=7B, # Frames=1fps2025.12 | 65.4 | |
| LongVU-7BModel Type=Open-Source 7B-8B Model2025.06 | 65.4 | |
| Video-XLLLM=7B, # Frames=1282025.12 | 64.9 | |
| GPT-4o2025.01 | 64.6 | |
| GPT-4oLLM=N/A, # Frames=N/A2025.12 | 64.6 | |
| GPT-4o(0513)Model Type=Closed-Source Model2025.06 | 64.6 | |
| GPT-4oAccess=Closed-Source2026.06 | 64.6 | |
| VideoMindModel Type=Agent Based Model2025.06 | 64.4 | |
| mPLUG-Owl3LLM=7B, # Frames=162025.12 | 63.7 | |
| VideoTreeType=Agentic video understanding framework2026.06 | 60.4 | |
| Tarsier-34B2025.01 | 58.2 | |
| VILA-1.5-40B2025.01 | 56.7 | |
| LongVALLM=7B, # Frames=2562025.12 | 56.3 | |
| Tarsier-7B2025.01 | 49.3 | |
| VideoChat2-7BModel Type=Open-Source 7B-8B Model2025.06 | 47.9 | |
| ShareGPT4Video-8BModel Type=Open-Source 7B-8B Model2025.06 | 46.4 | |
| VideoLLaMA-2-72BModel Type=Open-Source 72B-78B Model2025.06 | 45.6 |