Expert-level Multimodal Understanding on MMMU (Accuracy)
60.9AccuracyQwen3-VL-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-8BModel Category=Open-source General Models2025.11 | 60.9 | |
| PAPOModel Scale=7B, Decoding Strategy=greedy decoding2025.09 | 60.8 | |
| VAPOModel Scale=7B2025.09 | 60.2 | |
| VL-RethinkerModel Scale=7B2025.09 | 57.9 | |
| SenseNova-SIBase Architecture=Qwen3-VL-8B, Model Category=Ours2025.11 | 57.6 | |
| MM-EurekaModel Scale=7B2025.09 | 55.7 | |
| InternVL3-8BModel Category=Open-source General Models2025.11 | 55.6 | |
| Base ModelModel Scale=7B2025.09 | 52.7 | |
| VST-7B-SFTModel Category=Open-source SI Models2025.11 | 50.9 | |
| Bagel-7B-MoTModel Category=Open-source General Models2025.11 | 50.4 | |
| SenseNova-SIBase Architecture=Bagel-7B-MoT, Model Category=Ours2025.11 | 50.2 | |
| SenseNova-SIBase Architecture=InternVL3-8B, Model Category=Ours2025.11 | 49.4 | |
| Cambrian-S-7BModel Category=Open-source SI Models2025.11 | 47.1 | |
| SenseNova-SIBase Architecture=InternVL3-2B, Model Category=Ours2025.11 | 44.4 | |
| InternVL3-2BModel Category=Open-source General Models2025.11 | 43.2 |