General Multimodal Understanding on MMMU, MMBench, MME, ChartQA, AI2D, and HallBench Aggregate
72.6MMMU (Val)Gemini-2.0-Pro
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Gemini-2.0-ProModel Category=Closed-source2025.12 | 72.6 | 83 | 86.1 | 91.2 | 84.8 | 49.8 | 77.9 | |
| GPT-4o-latestModel Category=Closed-source2025.12 | 70.7 | 84.3 | 84.2 | 91.5 | 86.3 | 57 | 79 | |
| SR-MCR-7BModel Category=Our Models, Model size=7B, Separate Thinking Mode=false2025.12 | 67.6 | 91.2 | 87.3 | 94.5 | 88.2 | 59.4 | 81.4 | |
| SAIL-VL2-8B-ThinkingModel Category=Open-source, Model size=8B, Separate Thinking Mode=true2025.12 | 66.1 | 90.4 | 86 | 93.6 | 87.4 | 61.5 | 80.8 | |
| Keye-VL-8B-ThinkingModel Category=Open-source, Model size=8B, Separate Thinking Mode=true2025.12 | 63.4 | 81.7 | 83.5 | 88 | 86.4 | 62.7 | 77.6 | |
| Kimi-VL-A3B-ThinkingModel Category=Open-source, Model size=A3B, Separate Thinking Mode=true2025.12 | 60.4 | 89.7 | 87 | 92.1 | 83.1 | 58.3 | 78.4 | |
| InternVL3-8BModel Category=Open-source, Model size=8B2025.12 | 57.3 | 87.7 | 85.2 | 89.6 | 85.2 | 53.7 | 76.5 | |
| VL-Rethinker-7BModel Category=Open-source, Model size=7B2025.12 | 54.8 | 88.2 | 82.9 | 91.5 | 83.6 | 55.1 | 76 | |
| SR-MCR-3BModel Category=Our Models, Model size=3B, Separate Thinking Mode=false2025.12 | 52.8 | 86.9 | 80.8 | 90.9 | 84.3 | 51.9 | 74.6 | |
| VLAA-Thinker-7BModel Category=Open-source, Model size=7B2025.12 | 51.9 | 86.9 | 83.3 | 89.5 | 78.9 | 51.5 | 73.7 | |
| WeThink-7BModel Category=Open-source, Model size=7B2025.12 | 50.9 | 87.8 | 82.9 | 90.8 | 84.5 | 55.1 | 75.3 | |
| Qwen2.5-VL-7BModel Category=Open-source, Model size=7B2025.12 | 50.3 | 86.7 | 82.2 | 89.5 | 84 | 56 | 74.8 | |
| Qwen2.5-VL-3BModel Category=Open-source, Model size=3B2025.12 | 48.1 | 82.4 | 77.5 | 87 | 80.7 | 48.3 | 70.7 | |
| InternVL3-2BModel Category=Open-source, Model size=2B2025.12 | 47.1 | 84.3 | 77.4 | 80.4 | 78.7 | 41.4 | 68.2 |