Video Multimodal Understanding on VideoMMMU
79.4AccuracyGemini-2.5-Pro (minimal)
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5-Pro (minimal)Size=-2026.03 | 79.4 | |
| Seed-1.5-VLSize=20B2026.03 | 72.1 | |
| Pixelis (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 69.8 | |
| Qwen3-VL-30B-A3B-InstructSize=30B2026.03 | 68.7 | |
| Pixelis: SFT + RFTSize=8B2026.03 | 68.5 | |
| Pixelis: SFT + TTRLSize=8B2026.03 | 68.5 | |
| Pixelis: RFT + TTRLSize=8B2026.03 | 67.8 | |
| PRM (process reward; tools; 8B)Size=8B2026.03 | 67.6 | |
| Pixelis: SFT onlySize=8B2026.03 | 67.4 | |
| Late-fusionSize=8B2026.03 | 67 | |
| Step Self-Consistency (step-level)Size=8B2026.03 | 66.9 | |
| RA-TTA (retrieval-augmented)Size=8B2026.03 | 66.8 | |
| Pixel Reasoner (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 66.5 | |
| Pixelis: RFT onlySize=8B2026.03 | 66.5 | |
| Realistic TTA of VLMs (StatA)Size=8B2026.03 | 66.2 | |
| Pixelis: TTRL onlySize=8B2026.03 | 66.1 | |
| Qwen3-VL-8B-InstructSize=8B2026.03 | 65.3 | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 65.3 | |
| RV Self-Consistency (answer-only)Size=8B2026.03 | 65.1 | |
| GPT-5 (minimal)Size=-2026.03 | 61.6 | |
| GPT-4oCategory=Closed-source2026.01 | 61.2 | |
| Gemini-1.5-ProCategory=Closed-source2026.01 | 60.6 | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 54.9 | |
| Video-R1Training=SFT+ RL2026.03 | 52.4 | |
| NEO-ovArchitecture Type=Native, Parameter Scale=8B, Training Stage=Instruct2026.05 | 51.6 | |
| SDRLTraining=RL2026.03 | 51.3 | |
| SDRLTraining=RL, Training dataset=EventFlow2026.03 | 51.1 | |
| VideoRFTTraining=SFT+ RL, Input frames=16-frame2026.03 | 50.6 | |
| Qwen-VL-2.5-7B-OursCategory=Our Models2026.01 | 50 | |
| MiniCPM-V2.6-8BCategory=Open-source Base Models2026.01 | 49.8 | |
| Video-R1Training=RL2026.03 | 49.5 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=ours CoT2026.03 | 49.3 | |
| VideoRFTTraining=SFT, Input frames=16-frame2026.03 | 48.5 | |
| Qwen2.5-VL-7BTraining=None2026.03 | 48.4 | |
| Video-R1-7BCategory=Open-source Reasoning Models2026.01 | 48.1 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=video-r1 CoT2026.03 | 47.8 | |
| Video-R1Training=SFT2026.03 | 47.4 | |
| VideoRFTTraining=RL, Input frames=16-frame2026.03 | 47.4 | |
| Qwen-VL-2.5-7B-GRPOCategory=Our Models2026.01 | 47.3 | |
| Qwen-VL-2.5-7B-SFTCategory=Our Models2026.01 | 46 | |
| InternVL2.5-8BCategory=Open-source Base Models2026.01 | 44.2 | |
| R1-OneVision-7BCategory=Open-source Reasoning Models2026.01 | 44.1 | |
| Qwen-VL-2.5-7BCategory=Our Models2026.01 | 43.9 | |
| R1-VL-7BCategory=Open-source Reasoning Models2026.01 | 42.9 | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 42.7 | |
| NEO-ovArchitecture Type=Native, Parameter Scale=2B, Training Stage=Instruct2026.05 | 42.3 | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 41.9 | |
| Vision-R1-7BCategory=Open-source Reasoning Models2026.01 | 39.7 | |
| LLaVA-OneVision-7BTraining=None2026.03 | 33.8 | |
| LLaVA-OneVision-7BCategory=Open-source Video Models2026.01 | 31.2 | |
| LongVA-7BTraining=None2026.03 | 23.9 | |
| VILA-1.5-8BCategory=Open-source Video Models2026.01 | 20.8 | |
| VILA-1.5-8BTraining=None2026.03 | 20.8 |