Video Reasoning on MMMU Video
84.6AccuracyGPT-5-thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-thinkingModel Category=Closed-source Large-sized MLLM2025.11 | 84.6 | |
| Gemini 2.5 ProModel Category=Closed-source Large-sized MLLM2025.11 | 83.6 | |
| OpenAI O3Model Category=Closed-source Large-sized MLLM2025.11 | 83.3 | |
| Seed 1.5VLModel Category=Closed-source Large-sized MLLM2025.11 | 81.4 | |
| Qwen3-VL-235B-ThinkingModel Category=Open-source Large-sized MLLM2025.11 | 80 | |
| VideoChat-M1Model Category=Ours, Number of Parameters=37B2025.11 | 80 | |
| Qwen3-VL-235B-InstructModel Category=Open-source Large-sized MLLM2025.11 | 74.7 | |
| Qwen3-VL-8B-InstructModel Category=Open-source Medium-sized MLLM2025.11 | 65.3 | |
| GPT-4oModel Category=Closed-source Large-sized MLLM2025.11 | 61.2 | |
| Gemini 1.5 ProModel Category=Closed-source Large-sized MLLM2025.11 | 53.9 | |
| VideoChat-R1.5-7BModel Category=Open-source Medium-sized MLLM2025.11 | 51.4 | |
| Aria-28BModel Category=Open-source Large-sized MLLM2025.11 | 50.8 | |
| LLAVA-Video-72BModel Category=Open-source Large-sized MLLM2025.11 | 49.7 | |
| LLaVA-ov-72BModel Category=Open-source Large-sized MLLM2025.11 | 48.3 | |
| LLaVA-Video-7BModel Category=Open-source Medium-sized MLLM2025.11 | 36.1 | |
| LongVA-7BModel Category=Open-source Medium-sized MLLM2025.11 | 24 |