Video Question Answering on VideoMME Overall
73.6AccuracyVideo Panels (GPT-4.1-2025-04-14)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Video Panels (GPT-4.1-2025-04-14)#frames=32, Model Context Category=Commercial VLMs2025.09 | 73.6 | — | |
| GPT-4.1-2025-04-14#frames=32, Model Context Category=Commercial VLMs2025.09 | 72.5 | — | |
| Video Panels (LLaVA-Video 72B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 70.1 | — | |
| LLaVA-Video 72B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 69.8 | — | |
| EFlowFramework Type=Native Multi-turn Tool Invocation Video MLLM, Open-source status=Open-source2026.07 | 69.1 | — | |
| Qwen3-VL*Framework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 67.9 | — | |
| Video Panels (LLaVA-OV 72B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 67.7 | — | |
| Video Panels (Qwen-2.5VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 66.3 | — | |
| LLaVA-OV 72B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 66 | — | |
| Qwen-2.5VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 66 | — | |
| Rewatch-R1Framework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 65.6 | — | |
| VideoLLaMA 3 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 65.3 | — | |
| Video Panels (VideoLLaMA 3 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 65.3 | — | |
| Video-ZoomerFramework Type=Native Multi-turn Tool Invocation Video MLLM, Open-source status=Open-source2026.07 | 65.2 | — | |
| Video Panels (LLaVA-Video 7B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 64.4 | — | |
| LLaVA-Video 7B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 64.3 | — | |
| LongVTFramework Type=Native Multi-turn Tool Invocation Video MLLM, Open-source status=Open-source2026.07 | 64.3 | — | |
| VITALFramework Type=Native Multi-turn Tool Invocation Video MLLM, Open-source status=Open-source2026.07 | 64.1 | — | |
| Video Panels (Qwen-2.5VL 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 63.9 | — | |
| Video Panels (Qwen-2VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 63 | — | |
| Qwen2.5-VLFramework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 62.9 | — | |
| LLaVA-VideoFramework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 62.6 | — | |
| Qwen-2VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 62.4 | — | |
| Qwen-2.5VL 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 61.9 | — | |
| Video-R1Framework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 61.4 | — | |
| ConanFramework Type=Native Multi-turn Tool Invocation Video MLLM, Open-source status=Open-source2026.07 | 60.5 | — | |
| Video Panels (LLaVA-OV 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 58.9 | — | |
| LLaVA-OV 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 58.5 | — | |
| LLaVA-OneVisionFramework Type=Single-Turn Video MLLM, Open-source status=Open-source2026.07 | 56.2 | — | |
| Video Panels (GPT-4o-mini-2025-04-14)#frames=8, Model Context Category=Commercial VLMs2025.09 | 53.6 | — | |
| GPT-4o-mini-2025-04-14#frames=8, Model Context Category=Commercial VLMs2025.09 | 51.1 | — | |
| Video Panels (LLaVA-OV 0.5B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 44.3 | — | |
| LLaVA-OV 0.5B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 43.8 | — | |
| Video Panels (Video-LLaVA 7B)#frames=8, Model Context Category=Small-context VLMs2025.09 | 38.7 | — | |
| Video-LLaVA 7B#frames=8, Model Context Category=Small-context VLMs2025.09 | 37.1 | — | |
| BaselineModel=Video-LLaVA-7B, FLOPS Ratio=02025.03 | 30 | 4,120 | |
| TopVModel=Video-LLaVA-7B, FLOPS Ratio=51%2025.03 | 30 | 3,552 | |
| FastVModel=Video-LLaVA-7B, FLOPS Ratio=47%2025.03 | 29 | 6,711 | |
| Video Panels (VideoChat2-HD)#frames=16, Model Context Category=Small-context VLMs2025.09 | 25.4 | — | |
| VideoChat2-HD#frames=16, Model Context Category=Small-context VLMs2025.09 | 25.2 | — |