General Video Understanding on Video-MME (Accuracy)
87.8AccuracyGemini 1.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 1.5 ProSource=[60]2026.06 | 87.8 | |
| Gemini 1.5 ProSource=[59]2026.06 | 87.5 | |
| GPT-4o2026.06 | 86.9 | |
| Gemini 1.5 Flash2026.06 | 84.2 | |
| GPT-4o mini2026.06 | 82.3 | |
| Claude 3.5 Sonnet2026.06 | 80.5 | |
| SpatialClawBackbone=Gemma4-31B2026.06 | 77 | |
| Key-VL-1.5-8B2026.06 | 76.2 | |
| Molmo2-8B2026.06 | 75.8 | |
| Eagle-X5-7B2026.06 | 75.7 | |
| GLM-4V-9B2026.06 | 75.6 | |
| AdaCodecToken budget=comparable, Backbone=Qwen2-VL-7B2026.06 | 75.5 | |
| Qwen2-VL-7B2026.06 | 75.2 | |
| Gemini-1.5-ProFrames=322026.05 | 75 | |
| AdaCodecToken budget=1/7, Backbone=Qwen2-VL-7B2026.06 | 75 | |
| Gemini-1.5-ProSize=-, Zero-shot=true, Category=Proprietary Models2026.07 | 75 | |
| SpaceTools ToolshedBackbone=Gemma4-31B2026.06 | 74.6 | |
| Gemini-3.1-ProModel Type=Proprietary2026.05 | 74.3 | |
| MiniCPM-V-2.6-8B2026.06 | 73.5 | |
| SELECTSTREAM-Qwen3-VL-8BSize=8B, #Frames=1 fps, max 10242026.06 | 73.2 | |
| GPT-5.1Model Type=Proprietary2026.05 | 72.9 | |
| GPT-4oFrames=322026.05 | 71.9 | |
| GPT-4oSize=-, Zero-shot=true, Category=Proprietary Models2026.07 | 71.9 | |
| Qwen3-VL-8BSize=8B, #Frames=2 fps, max 20482026.06 | 71.4 | |
| LLaVA-Video-7B2026.06 | 69.7 | |
| VideoChat-Flash-7B2026.06 | 69.7 | |
| CoPE-VideoLM-7B2026.06 | 69.4 | |
| Molmo2-O-7B2026.06 | 69.2 | |
| InternVL2-8B2026.06 | 68.6 | |
| Streamo-7BSize=7B, #Frames=1 fps2026.06 | 67.9 | |
| SELECTSTREAM-Qwen2.5-VL-7BSize=7B, #Frames=1 fps, max 10242026.06 | 67.8 | |
| Qwen-2.5-VLSize=7B, Zero-shot=true, Category=Supervised Fine-Tuning Models2026.07 | 66.1 | |
| InternVL3-8BModel Type=Open-Source2026.05 | 65.6 | |
| Qwen3-VL-8B + CRPOBackbone=Qwen3-VL-8B, Post-training Method=CRPO2026.05 | 65.6 | |
| TimeThinkSize=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 65.5 | |
| PLM-7B2026.06 | 65.4 | |
| VideoChat-R1.5Size=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 65.2 | |
| Qwen3-VL-8B + ArrowRLBackbone=Qwen3-VL-8B, Post-training Method=ArrowRL2026.05 | 65.1 | |
| Qwen2.5-VL-7BSize=7B, #Frames=max 7682026.06 | 65.1 | |
| Qwen3-VL-8B + T-GRPOBackbone=Qwen3-VL-8B, Post-training Method=T-GRPO2026.05 | 65 | |
| Qwen3-VL-8BBackbone=Qwen3-VL-8B, Post-training Method=base2026.05 | 64.9 | |
| Qwen3-VL-8B + GRPOBackbone=Qwen3-VL-8B, Post-training Method=GRPO2026.05 | 64.5 | |
| ReMoRa-7B2026.06 | 64.4 | |
| Qwen2.5-VL-GRPOSize=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 64.3 | |
| InternVL2.5Size=8B, Zero-shot=true, Category=Supervised Fine-Tuning Models2026.07 | 64.2 | |
| VideoChat-R1Size=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 64.1 | |
| VideoLatent-7B (Ours)Frames=64, Architecture Category=Latent MLLM2026.06 | 63.8 | |
| Qwen2.5-VL-SFTSize=7B, Zero-shot=true, Category=Supervised Fine-Tuning Models2026.07 | 63.6 | |
| Qwen3-VL-4B + T-GRPOBackbone=Qwen3-VL-4B, Post-training Method=T-GRPO2026.05 | 63.4 | |
| Qwen2-VL-7BSize=7B, #Frames=642026.06 | 63.3 | |
| LLaVA-Video-7BSize=7B, #Frames=642026.06 | 63.3 | |
| Qwen3-VL-4B + CRPOBackbone=Qwen3-VL-4B, Post-training Method=CRPO2026.05 | 63 | |
| Qwen2.5-VL-7BFrames=64, Architecture Category=Standard MLLM, CoT=false2026.06 | 62.7 | |
| Qwen3-VL-4BBackbone=Qwen3-VL-4B, Post-training Method=base2026.05 | 62.2 | |
| Qwen3-VL-4B + ArrowRLBackbone=Qwen3-VL-4B, Post-training Method=ArrowRL2026.05 | 62.2 | |
| Qwen3-VL-4B + GRPOBackbone=Qwen3-VL-4B, Post-training Method=GRPO2026.05 | 62.1 | |
| Mull-Token-7BFrames=64, Architecture Category=Latent MLLM2026.06 | 62.1 | |
| LVR-7BFrames=64, Architecture Category=Latent MLLM2026.06 | 62 | |
| StreamForest-7B (FT-drive)Size=7B, #Frames=1 fps2026.06 | 61.9 | |
| StreamForest-7BSize=7B, #Frames=1 fps2026.06 | 61.4 | |
| VideoLatent-7B (Ours)Frames=32, Architecture Category=Latent MLLM2026.06 | 61.4 | |
| Video-R1-7BFrames=64, Architecture Category=Standard MLLM2026.06 | 61.4 | |
| Open-o3-Video-7BFrames=64, Architecture Category=Standard MLLM2026.06 | 61.4 | |
| Video-R1Size=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 61.4 | |
| Open-o3-Video-7BFrames=32, Architecture Category=Standard MLLM2026.06 | 61.3 | |
| STORMFrames=322026.05 | 61 | |
| Qwen2.5-VL-7B-Instruct(CoT)Frames=322026.05 | 60.8 | |
| LongVU-7BSize=7B, #Frames=1 fps2026.06 | 60.6 | |
| LVR-7BFrames=32, Architecture Category=Latent MLLM2026.06 | 60.5 | |
| Qwen2.5-VL-7BModel Type=Open-Source2026.05 | 60.3 | |
| LLaVA-Next-VideoFrames=322026.05 | 60.2 | |
| ArrowRL*Model Type=Open-Source2026.05 | 60.1 | |
| VILA-1.5Size=40B, Zero-shot=true, Category=Supervised Fine-Tuning Models2026.07 | 60.1 | |
| Video-Thinker-7BFrames=32, method_variant=think with video2026.05 | 60 | |
| VideoRFT-7BFrames=32, Architecture Category=Standard MLLM2026.06 | 59.8 | |
| VideoRFTSize=7B, Zero-shot=true, Category=Reinforcement Fine-Tuning Models2026.07 | 59.8 | |
| LongVT-RFTFrames=32, method_variant=think with tools2026.05 | 59.7 | |
| Qwen2.5-VL-7B (CoT)Frames=64, Architecture Category=Standard MLLM, CoT=true2026.06 | 59.6 | |
| Mull-Token-7BFrames=32, Architecture Category=Latent MLLM2026.06 | 59.5 | |
| GPT-4VSize=-, Zero-shot=true, Category=Proprietary Models2026.07 | 59.5 | |
| Qwen2.5-VL-7BFrames=32, Architecture Category=Standard MLLM, CoT=false2026.06 | 59.4 | |
| VideoLatent-7B (Ours)Frames=16, Architecture Category=Latent MLLM2026.06 | 59.3 | |
| Video-R1-7BFrames=32, Architecture Category=Standard MLLM2026.06 | 59.3 | |
| LVR-7BFrames=16, Architecture Category=Latent MLLM2026.06 | 59.1 | |
| LLaVA-OneVision-7BSize=7B, #Frames=322026.06 | 58.2 | |
| LLaVA-OVSize=7B, Zero-shot=true, Category=Supervised Fine-Tuning Models2026.07 | 58.2 | |
| LLaVA-OV-7BModel Type=Open-Source2026.05 | 57.6 | |
| Video-R1-7BFrames=322026.05 | 57.4 | |
| Video-R1-7BFrames=16, Architecture Category=Standard MLLM2026.06 | 57.4 | |
| Dispider-7BSize=7B, #Frames=1 fps2026.06 | 57.2 | |
| Mull-Token-7BFrames=16, Architecture Category=Latent MLLM2026.06 | 57 | |
| Qwen2.5-VL-7BFrames=16, Architecture Category=Standard MLLM, CoT=false2026.06 | 56.6 | |
| Qwen2.5-VL-7B (CoT)Frames=32, Architecture Category=Standard MLLM, CoT=true2026.06 | 56.6 | |
| KangarooFrames=322026.05 | 56 | |
| VideoRefer2026.05 | 55.9 | |
| SWIM2026.05 | 55.9 | |
| LLaVA-Octopus2026.05 | 55.7 | |
| Qwen2.5-VL-7B-SFTFrames=322026.05 | 55.4 | |
| VideoLLaMA2.12026.05 | 54.9 | |
| VideoChat2-HD2026.05 | 54.6 |