Long-form Video Understanding on LongVideoBench (Accuracy)
82.3AccuracyVideoChat-M1
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoChat-M1Model Category=Ours, Number of Parameters=37B2025.11 | 82.3 | |
| GPT-4oAgent=disabled2026.05 | 81.3 | |
| Gemini 2.5 ProModel Category=Closed-source Large-sized MLLM2025.11 | 78.7 | |
| Seed1.5-VLThinking mode=false, Evaluation frame rate=1 FPS2025.05 | 74.4 | |
| Seed1.5-VLThinking mode=true, Evaluation frame rate=1 FPS2025.05 | 74 | |
| Seed 1.5VLModel Category=Closed-source Large-sized MLLM2025.11 | 74 | |
| DeepVideoDiscoveryModel Category=Agent-based Methods2025.11 | 71.6 | |
| NVILAModel Size=8B, Video Tokens=8K2025.11 | 70.1 | |
| InternVL3.5-241B-A28B2026.01 | 67.1 | |
| InternVL-3.5-241BModel Category=Open-source Large-sized MLLM2025.11 | 67.1 | |
| GPT-4o2024.09 | 66.7 | |
| GPT-4oEvaluation frame rate=1 FPS2025.05 | 66.7 | |
| GPT-4oSize=-2026.02 | 66.7 | |
| GPT-4oCategory=Proprietary Models2025.12 | 66.7 | |
| GPT-4oModel Category=Proprietary Models2025.12 | 66.7 | |
| GPT-4oModel Category=Closed-source Large-sized MLLM2025.11 | 66.7 | |
| GPT4o2025.11 | 66.7 | |
| VideoSeeker-8BAgent=enabled2026.05 | 66.5 | |
| Eagle-2.5-8BModel Category=Open-source Medium-sized MLLM2025.11 | 66.4 | |
| InternVL-3-78BModel Category=Open-source Large-sized MLLM2025.11 | 65.7 | |
| InternVL-3.5-78BModel Category=Open-source Large-sized MLLM2025.11 | 65.7 | |
| VideoChat-A1Model Category=Agent-based Methods2025.11 | 65.4 | |
| VideoRAG-72BModel Category=Agent-based Methods2025.11 | 65.4 | |
| LLAVA-Video-72BModel Category=Open-source Large-sized MLLM2025.11 | 64.9 | |
| VideoChat-Flash-7BModel Category=Open-source Medium-sized MLLM2025.11 | 64.7 | |
| Qwen3-VL-8BAgent=disabled2026.05 | 64.6 | |
| Gemini-1.5-Pro2024.09 | 64.4 | |
| Aria-28BModel Category=Open-source Large-sized MLLM2025.11 | 64.2 | |
| VideoSeeker-4BAgent=enabled2026.05 | 64.2 | |
| Gemini-1.5-ProSize=-2026.02 | 64 | |
| Gemini-1.5-proCategory=Proprietary Models2025.12 | 64 | |
| Gemini-1.5-proModel Category=Proprietary Models2025.12 | 64 | |
| Gemini 1.5 ProModel Category=Closed-source Large-sized MLLM2025.11 | 64 | |
| Gemini-1.5-Pro2025.11 | 64 | |
| Qwen3-VL-8B-Thinking + MVPFrames=1282026.01 | 63.8 | |
| InternVL-2.5-78BModel Category=Open-source Large-sized MLLM2025.11 | 63.6 | |
| Qwen3-VL-8B-ThinkingFrames=1282026.01 | 62.9 | |
| VideoChat-R1.5-7BModel Category=Open-source Medium-sized MLLM2025.11 | 62.6 | |
| Qwen3-VL-4BAgent=disabled2026.05 | 62.6 | |
| OryxSize=34B2024.09 | 62.2 | |
| InternVL3.5Size=8B2026.02 | 62.1 | |
| InternVL-3.5-8BModel Category=Open-source Medium-sized MLLM2025.11 | 62.1 | |
| Oryx-1.5Size=32B2024.09 | 62 | |
| Qwen3-VL-8B-Thinking + MVPFrames=642026.01 | 62 | |
| LLaVA-VideoSize=72B2026.02 | 61.9 | |
| Demo-ICL (SFT)Size=7B2026.02 | 61.6 | |
| VidLaDA#S (Model Size)=8B, #F (Number of input frames)=64, Architecture=DLM Baseline2026.01 | 61.4 | |
| LLaVA-OneVisionSize=72B2024.09 | 61.3 | |
| LLaVA-OneVisionSize=72B2026.02 | 61.3 | |
| LLaVA-ov-72BModel Category=Open-source Large-sized MLLM2025.11 | 61.3 | |
| Demo-ICLSize=7B2026.02 | 61.2 | |
| InternVL3.5-8B + MVPFrames=642026.01 | 61 | |
| VideoXL2-8BModel Category=Open-source Medium-sized MLLM2025.11 | 61 | |
| GPT-4V2024.09 | 60.7 | |
| Qwen2.5-VLSize=72B2026.02 | 60.7 | |
| Ola-Video (Base)Size=7B2026.02 | 60.7 | |
| Qwen2.5-VL-72BFrames=7682026.01 | 60.7 | |
| Qwen2.5-VL-72BModel Category=Open-source Large-sized MLLM2025.11 | 60.7 | |
| InternVideo2.5-7BModel Category=Open-source Medium-sized MLLM2025.11 | 60.6 | |
| Hour-LLaVAModel Size=7B2025.11 | 60.4 | |
| Qwen2.5-VL#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline, reproduced=true2026.01 | 60.2 | |
| InternVL2.5#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline2026.01 | 60 | |
| InternVL-2.5-8BModel Category=Open-source Medium-sized MLLM2025.11 | 60 | |
| VideoLLaMA3Size=7B2026.02 | 59.8 | |
| Qwen3-VL-8B-ThinkingFrames=642026.01 | 59.6 | |
| InternVL3.5-8B + MVPFrames=1282026.01 | 59.5 | |
| Streamo-7BCategory=Streamo Framework, Model Size=7B, Base Model=Qwen2.5-VL-7B2025.12 | 59.2 | |
| Streamo-7BModel Category=Streamo Framework, Parameters=7B2025.12 | 59.2 | |
| StreamingVLM-7BCategory=Open-source Online Models, Model Size=7B2025.12 | 59 | |
| StreamingVLM-7BModel Category=Open-source Online Models, Parameters=7B2025.12 | 59 | |
| InternVL-3-8BModel Category=Open-source Medium-sized MLLM2025.11 | 58.8 | |
| LLaDA-V#S (Model Size)=8B, #F (Number of input frames)=32, Architecture=DLM Baseline, reproduced=true2026.01 | 58.6 | |
| Qwen2.5-VL-7B-Instruct + MVPFrames=1282026.01 | 58.6 | |
| GPT-4oParams=-, Frame=322026.02 | 58.5 | |
| LLaVA-VideoSize=7B2026.02 | 58.2 | |
| LLaVA-Video#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline, reproduced=true2026.01 | 58.2 | |
| LLaVA-Video-7BModel Category=Open-source Medium-sized MLLM2025.11 | 58.2 | |
| LongRL-7BModel Category=Open-source Medium-sized MLLM2025.11 | 58.1 | |
| LongVILA-R1Model Size=7B2025.11 | 58 | |
| REVISOR (Ours)Model Size=7B, Video Tokens=8K2025.11 | 57.5 | |
| Qwen2.5-VL*Model Size=7B, Video Tokens=8K, Training=text-based reflection mechanism2025.11 | 57.4 | |
| LongVILA-7BModel Category=Open-source Medium-sized MLLM2025.11 | 57.1 | |
| LLaVA-OneVision#S (Model Size)=7B, #F (Number of input frames)=32, Architecture=AR Baseline, reproduced=true2026.01 | 56.5 | |
| Streamo-2BCategory=Streamo Framework, Model Size=2B, Base Model=InternVL3-2B2025.12 | 56.5 | |
| Qwen2.5-VL⋆Model Size=7B, Video Tokens=8K, Status=reproduction2025.11 | 56.5 | |
| LLaVA-OneVisionModel Size=7B, Video Tokens=6K2025.11 | 56.4 | |
| VL-RethinkerModel Size=7B, Video Tokens=8K2025.11 | 56.4 | |
| Oryx-1.5Size=7B2024.09 | 56.3 | |
| LLaVA-OneVisionSize=7B2026.02 | 56.3 | |
| Jigsaw-7BFrames=1282026.01 | 56.2 | |
| Streamo-3BCategory=Streamo Framework, Model Size=3B, Base Model=Qwen2.5-VL-3B2025.12 | 56.2 | |
| Streamo-3BModel Category=Streamo Framework, Parameters=3B2025.12 | 56.2 | |
| Streamo-4BCategory=Streamo Framework, Model Size=4B, Base Model=Qwen3-VL-4B2025.12 | 56.1 | |
| Qwen2.5-VLSize=7B2026.02 | 56 | |
| Qwen2.5-VL-7BCategory=Streamo Framework, Model Size=7B2025.12 | 56 | |
| Qwen2.5-VL-7BModel Category=Streamo Framework, Parameters=7B2025.12 | 56 | |
| Qwen2.5-VL-7BModel Category=Open-source Medium-sized MLLM2025.11 | 56 | |
| Qwen2.5-VL-7B-Instruct + MVPFrames=642026.01 | 55.9 | |
| VambaModel Size=10B2025.11 | 55.9 | |
| InternVL3.5-8BFrames=1282026.01 | 55.6 |