Multimodal Understanding on MMVU
75.4AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oCategory=Closed-source2026.01 | 75.4 | |
| Qwen-VL-2.5-7B-OursCategory=Our Models2026.01 | 65.6 | |
| Qwen-VL-2.5-7B-SFTCategory=Our Models2026.01 | 62.7 | |
| Video-R1-7BCategory=Open-source Reasoning Models2026.01 | 62.5 | |
| Qwen-VL-2.5-7B-GRPOCategory=Our Models2026.01 | 62.1 | |
| CogniRoute2026.06 | 60.7 | |
| Qwen3 Omni 30BParameters=30B2026.06 | 59.8 | |
| R1-VL-7BCategory=Open-source Reasoning Models2026.01 | 59.7 | |
| Vision-R1-7BCategory=Open-source Reasoning Models2026.01 | 57.6 | |
| Qwen-VL-2.5-7BCategory=Our Models2026.01 | 57.6 | |
| R1-OneVision-7BCategory=Open-source Reasoning Models2026.01 | 55.2 | |
| InternVL2.5-8BCategory=Open-source Base Models2026.01 | 54.9 | |
| MiniCPM-V2.6-8BCategory=Open-source Base Models2026.01 | 52.4 | |
| LLaVA-OneVision-7BCategory=Open-source Video Models2026.01 | 49.2 | |
| Penguin-VLParameters=2B2026.03 | 42.7 | |
| InternVL3.5Parameters=2B2026.03 | 42.7 | |
| Qwen3-VLParameters=2B2026.03 | 41.7 | |
| Gemma3n E2B-itParameters=E2B-it, Inference template=Penguin's template, Frame budget=322026.03 | 34.5 | |
| SmolVLM2Parameters=2.2B2026.03 | 33.5 | |
| VILA-1.5-8BCategory=Open-source Video Models2026.01 | 31.5 |