Video Understanding on VideoMME w/o sub. Overall
75Overall AccuracyGemini-1.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-1.5-ProCategory=Proprietary Models, Size=-2026.06 | 75 | |
| GPT-4oCategory=Proprietary Models, Size=-2026.06 | 71.9 | |
| OmniAgentCategory=Ours, Size=7B, incorporate audio signals=true2026.06 | 67.8 | |
| Zoom-ZeroCategory=Open-Source Agentic Models, Size=7B2026.06 | 66 | |
| LongVILA-R1Category=Open-Source Thinking Models, Size=7B2026.06 | 65.1 | |
| Qwen2.5-VLCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 65.1 | |
| Qwen2.5-OmniCategory=Open-Source Non-Thinking Models, Size=7B, incorporate audio signals=true2026.06 | 64.8 | |
| VITALCategory=Open-Source Agentic Models, Size=7B2026.06 | 64.1 | |
| Open-o3 VideoCategory=Open-Source Thinking Models, Size=7B2026.06 | 63.6 | |
| Video-R1Category=Open-Source Thinking Models, Size=7B2026.06 | 61.4 | |
| LongVUCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 60.6 | |
| LongVILACategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 60.1 | |
| VideoRFTCategory=Open-Source Thinking Models, Size=7B2026.06 | 59.8 | |
| Video-CoMCategory=Open-Source Agentic Models, Size=7B2026.06 | 59.4 | |
| LLaVA-OneVisionCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 58.2 | |
| VambaCategory=Open-Source Non-Thinking Models, Size=10B2026.06 | 57.8 | |
| KangarooCategory=Open-Source Non-Thinking Models, Size=8B2026.06 | 56 | |
| LongVTCategory=Open-Source Agentic Models, Size=7B2026.06 | 55.9 | |
| VISTACategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 55.5 | |
| LongVACategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 51.8 | |
| VideoLLaMA2Category=Open-Source Non-Thinking Models, Size=7B, incorporate audio signals=true2026.06 | 47.9 | |
| VideoChat-TCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 46.3 |