Multimodal Reasoning on MMMU-Pro (Pass@1)
70.64Pass@1GPT-5-Nano-High
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-Nano-HighModel Category=Closed-source Models2026.02 | 70.64 | |
| Qwen3-VL-8B-ThinkingSeries=Qwen3-VL-8B2026.02 | 70.29 | |
| Qwen3-VL-8B-DeepVisionSeries=Qwen3-VL-8B, Training Algorithm=GSPO2026.02 | 70.29 | |
| MiMo-VL-7B-DeepVisionSeries=MiMo-VL-7B, Training Algorithm=GSPO2026.02 | 69.19 | |
| Gemini 1.5 Prothinking=true, decoding=greedy2025.05 | 68.8 | |
| Qwen3-VL-8B-InstructSeries=Qwen3-VL-8B2026.02 | 67.69 | |
| Seed 1.5-VLthinking=true, decoding=greedy2025.05 | 67.6 | |
| MiMo-VL-7B-OpenMMReasonerSeries=MiMo-VL-7B2026.02 | 66.82 | |
| OpenAI o1thinking=true, decoding=greedy2025.05 | 66.4 | |
| MiMo-VL-7B-MM-EurekaSeries=MiMo-VL-7B2026.02 | 65.78 | |
| Gemini-2.5-Flash-LiteModel Category=Closed-source Models2026.02 | 65.08 | |
| MiMo-VL-7B-RL-2508Series=MiMo-VL-7B2026.02 | 63.87 | |
| MiMo-VL-7B-MathBookSeries=MiMo-VL-7B2026.02 | 63.47 | |
| MiMo-VL-7B-SFT-2508Series=MiMo-VL-7B2026.02 | 60.69 | |
| Seed 1.5-VLthinking=false, decoding=greedy2025.05 | 59.9 | |
| GPT-4othinking=false, decoding=greedy2025.05 | 54.5 | |
| Qwen 2.5-VL 72Bthinking=false, decoding=greedy2025.05 | 51.1 | |
| Claude 3.7 Sonnetthinking=true, decoding=sampling2025.05 | 50.1 |