Video Question Answering on VideoMME Medium
72.9AccuracyVideo Panels (GPT-4.1-2025-04-14)
Evaluation Results
| Method | Links | |
|---|---|---|
| Video Panels (GPT-4.1-2025-04-14)#frames=32, Model Context Category=Commercial VLMs2025.09 | 72.9 | |
| GPT-4.1-2025-04-14#frames=32, Model Context Category=Commercial VLMs2025.09 | 68.9 | |
| Video Panels (LLaVA-Video 72B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 68.6 | |
| LLaVA-Video 72B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 67.8 | |
| Qwen-2.5VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 67.6 | |
| Video Panels (Qwen-2.5VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 66.9 | |
| Video Panels (LLaVA-OV 72B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 66.4 | |
| VideoLLaMA 3 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 64.6 | |
| Video Panels (Qwen-2.5VL 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 64 | |
| Video Panels (VideoLLaMA 3 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 63.7 | |
| LLaVA-OV 72B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 62.9 | |
| Video Panels (Qwen-2VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 62.9 | |
| Qwen-2VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 62.7 | |
| LLaVA-Video 7B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 62.3 | |
| Video Panels (LLaVA-Video 7B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 62.2 | |
| GPT-4O + VSLSSearching Modality=Unimodal, Frame=322025.08 | 61.9 | |
| GPT-4O + VSISearching Modality=Multimodal, Frame=322025.08 | 61.7 | |
| GPT-4O + TstarSearching Modality=Unimodal, Frame=322025.08 | 61.6 | |
| Qwen-2.5VL 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 61.6 | |
| GPT-4OSearching Modality=N/A, Frame=82025.08 | 61.2 | |
| GPT-4O + VSLSSearching Modality=Unimodal, Frame=82025.08 | 61.1 | |
| GPT-4OSearching Modality=N/A, Frame=322025.08 | 61 | |
| GPT-4O + TstarSearching Modality=Unimodal, Frame=82025.08 | 59.7 | |
| GPT-4O + VSISearching Modality=Multimodal, Frame=82025.08 | 59.5 | |
| LLaVA-OV 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 56.7 | |
| Video Panels (LLaVA-OV 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 56.2 | |
| Video Panels (GPT-4o-mini-2025-04-14)#frames=8, Model Context Category=Commercial VLMs2025.09 | 53 | |
| QWEN2.5-VL-7B-INSTRUCT+TstarSearching Modality=Unimodal, Frame=82025.08 | 50 | |
| QWEN2.5-VL-7B-INSTRUCT+VSLSSearching Modality=Unimodal, Frame=322025.08 | 50 | |
| GPT-4o-mini-2025-04-14#frames=8, Model Context Category=Commercial VLMs2025.09 | 50 | |
| QWEN2.5-VL-7B-INSTRUCT+VSLSSearching Modality=Unimodal, Frame=82025.08 | 49.6 | |
| QWEN2.5-VL-7B-INSTRUCT+VSISearching Modality=Multimodal, Frame=82025.08 | 48.5 | |
| QWEN2.5-VL-7B-INSTRUCTSearching Modality=N/A, Frame=82025.08 | 47.3 | |
| QWEN2.5-VL-7B-INSTRUCT+VSISearching Modality=Multimodal, Frame=322025.08 | 45.8 | |
| QWEN2.5-VL-7B-INSTRUCT+TstarSearching Modality=Unimodal, Frame=322025.08 | 45.6 | |
| Video Panels (LLaVA-OV 0.5B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 42.6 | |
| LLaVA-Video-7B-Qwen2+VSISearching Modality=Multimodal, Frame=82025.08 | 42.2 | |
| LLaVA-Video-7B-Qwen2+VSISearching Modality=Multimodal, Frame=322025.08 | 42.2 | |
| LLaVA-OV 0.5B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 40.9 | |
| LLaVA-Video-7B-Qwen2+TstarSearching Modality=Unimodal, Frame=82025.08 | 40.4 | |
| QWEN2.5-VL-7B-INSTRUCTSearching Modality=N/A, Frame=322025.08 | 39.9 | |
| LLaVA-Video-7B-Qwen2Searching Modality=N/A, Frame=82025.08 | 39.7 | |
| LLaVA-Video-7B-Qwen2+TstarSearching Modality=Unimodal, Frame=322025.08 | 39.6 | |
| LLaVA-Video-7B-Qwen2+VSLSSearching Modality=Unimodal, Frame=322025.08 | 39 | |
| LLaVA-Video-7B-Qwen2+VSLSSearching Modality=Unimodal, Frame=82025.08 | 38.5 | |
| Video-LLaVAVision Size (M)=4252024.12 | 38.1 | |
| Video-PandaVision Size (M)=452024.12 | 37.9 | |
| Video Panels (Video-LLaVA 7B)#frames=8, Model Context Category=Small-context VLMs2025.09 | 37.9 | |
| Video-LLaVA 7B#frames=8, Model Context Category=Small-context VLMs2025.09 | 36.6 | |
| LLaVA-Video-7B-Qwen2Searching Modality=N/A, Frame=322025.08 | 36.5 | |
| Video-ChatGPTVision Size (M)=3072024.12 | 36 | |
| Video Panels (VideoChat2-HD)#frames=16, Model Context Category=Small-context VLMs2025.09 | 26.7 | |
| VideoChat2-HD#frames=16, Model Context Category=Small-context VLMs2025.09 | 26.4 |