Multimodal Benchmarking on MMMU
78.1AccuracyQwen3-VL-32B-Thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-32B-ThinkingParameters=32B, Thinking Mode=true2026.01 | 78.1 | |
| Qwen3-VL-8B-ThinkingParameters=8B, Thinking Mode=true2026.01 | 74.1 | |
| EvoCUA-32BParameters=32B2026.01 | 68.11 | |
| EvoCUA-8BParameters=8B2026.01 | 62.11 | |
| OpenCUA-72BParameters=72B2026.01 | 60.67 | |
| EvoCUA-OpenCUA-72BBase Model=OpenCUA-72B2026.01 | 59.22 | |
| INF-LLaVA*Source=Ours, larger_dataset=true2024.07 | 37.2 | |
| INF-LLaVASource=Ours2024.07 | 37 | |
| LLaVA1.5Source=CVPR’242024.07 | 36.4 | |
| ConvLLaVASource=arXiv’242024.07 | 35.8 | |
| DeepStack-L-HDSource=arXiv’242024.07 | 35.6 | |
| AnyGPTSource=arXiv’242024.07 | 30.6 | |
| GenLLaVASource=arXiv’242024.07 | 29.7 | |
| GILLSource=NeurIPS’232024.07 | 28.8 | |
| MGIESource=ICLR’242024.07 | 25.6 |