Multi-discipline Multimodal Understanding on MMMU
84.2AccuracyGPT-5
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-52025.12 | 84.2 | — | — | — | — | — | |
| GPT-5-Mini2025.12 | 79 | — | — | — | — | — | |
| InternVL3.5-38B2025.12 | 76.9 | — | — | — | — | — | |
| InternVL3.5-38B2025.06 | 76.9 | 82.8 | — | — | — | — | |
| GenRecal (InternVL3.5-8B)Teacher VLM=InternVL3.5-38B2025.06 | 76.5 | 83.4 | — | — | — | — | |
| GenRecal (InternVL3.5-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 76.3 | 83.4 | — | — | — | — | |
| Qwen3-VL-32B2025.12 | 76 | — | — | — | — | — | |
| Qwen3-VL-32B2025.06 | 76 | 82.5 | — | — | — | — | |
| GLM-4.5V2025.12 | 75.4 | — | — | — | — | — | |
| Gemini-2.5-Pro2025.12 | 74.7 | — | — | — | — | — | |
| Claude-4-Sonnet2025.12 | 74.4 | — | — | — | — | — | |
| InternVL3-8B-MastersSize=8B2025.12 | 74 | — | — | — | — | — | |
| GPT-4.12025.12 | 74 | — | — | — | — | — | |
| MastersBase Model=InternVL3-8B2025.12 | 74 | — | — | — | — | — | |
| InternVL3.5-8BSize=8B2025.12 | 73.4 | — | — | — | — | — | |
| InternVL3.5-8B2025.06 | 73.4 | 80.4 | — | — | — | — | |
| Qwen3-VL-8B-MastersSize=8B2025.12 | 72.9 | — | — | — | — | — | |
| MastersBase Model=Qwen3-VL-8B2025.12 | 72.9 | — | — | — | — | — | |
| InternVL3.5-8B-MastersSize=8B2025.12 | 72.7 | — | — | — | — | — | |
| MastersBase Model=InternVL3.5-8B2025.12 | 72.7 | — | — | — | — | — | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 72.7 | 80.1 | — | — | — | — | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=InternVL3.5-38B2025.06 | 72.5 | 80 | — | — | — | — | |
| InternVL3-78B2025.12 | 72.2 | — | — | — | — | — | |
| Keye-VL-8BSize=8B2025.12 | 71.4 | — | — | — | — | — | |
| Keye-VL-1.5-8BSize=8B2025.12 | 71.4 | — | — | — | — | — | |
| Qwen2.5-VL-7B-MastersSize=7B2025.12 | 71.3 | — | — | — | — | — | |
| MastersBase Model=Qwen2.5-VL-7B2025.12 | 71.3 | — | — | — | — | — | |
| Claude-3.7-Sonnet2025.12 | 71 | — | — | — | — | — | |
| Qwen3-VL-4B-MastersSize=4B2025.12 | 70.3 | — | — | — | — | — | |
| Gemini-2.0-Flash2025.12 | 69.9 | — | — | — | — | — | |
| Qwen3-VL-8BSize=8B2025.12 | 69.6 | — | — | — | — | — | |
| Qwen3-VL-8B2025.06 | 69.6 | 77 | — | — | — | — | |
| VLsI-7BParameters=7B2024.12 | 69.3 | — | — | — | — | — | |
| GPT-4o2024.09 | 69.2 | — | — | — | — | — | |
| GPT-4o2025.09 | 69.2 | — | — | — | — | — | |
| GPT-4o2025.12 | 69.1 | — | — | — | — | — | |
| Claude-3.5-Sonnet2025.12 | 68.3 | — | — | — | — | — | |
| Qwen2.5-VL-72B2025.12 | 68.2 | — | — | — | — | — | |
| GLM-4.1V-9BSize=9B2025.12 | 68 | — | — | — | — | — | |
| Qwen3-VL-4BSize=4B2025.12 | 67.4 | — | — | — | — | — | |
| MiMo-VL-8BSize=8B2025.12 | 66.7 | — | — | — | — | — | |
| InternVL3.5-4BSize=4B2025.12 | 66.6 | — | — | — | — | — | |
| InternVL3.5-2B-MastersSize=2B2025.12 | 64.6 | — | — | — | — | — | |
| InternVL3.5-4B-MastersSize=4B2025.12 | 63.2 | — | — | — | — | — | |
| GPT-4oModel Category=Proprietary VLMs2024.06 | 62.8 | — | — | — | — | — | |
| SketchThinker-R1-7BMethod Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 62.8 | — | — | 64.3 | 97.7 | — | |
| InternVL3-8BSize=8B2025.12 | 62.7 | — | — | — | — | — | |
| Gemini-1.5-Pro2025.12 | 62.2 | — | — | — | — | — | |
| Gemini-1.5-Pro2025.09 | 62.2 | — | — | — | — | — | |
| Qwen2.5-VL-3B-MastersSize=3B2025.12 | 61.2 | — | — | — | — | — | |
| Vanilla-R1Method Category=Direct Inference, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 61 | — | — | 182.2 | 33.5 | — | |
| VeriThinkerMethod Category=Supervised-fine-tuning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 60.1 | — | — | 105.8 | 56.8 | — | |
| EMOVAModel Size=72B2024.09 | 59.7 | — | — | — | — | — | |
| NVLM-72B2025.12 | 59.7 | — | — | — | — | — | |
| InternVL3-2B-MastersSize=2B2025.12 | 59.6 | — | — | — | — | — | |
| L1Method Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 59.5 | — | — | 136.8 | 43.5 | — | |
| C3oTMethod Category=Supervised-fine-tuning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 59.3 | — | — | 127.1 | 46.7 | — | |
| ThinkPruneMethod Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 59.2 | — | — | 104.9 | 56.4 | — | |
| InternVL3.5-2BSize=2B2025.12 | 59 | — | — | — | — | — | |
| Chain-of-DraftMethod Category=Prompt-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 58.9 | — | — | 86.3 | 68.3 | — | |
| ETTCAggregation Strategy=ETTC, Model Scale=7B-12B Ensemble2026.05 | 58.63 | — | — | — | — | — | |
| Constrained CoTMethod Category=Prompt-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 58.6 | — | — | 78.2 | 74.9 | — | |
| Qwen2.5-VLUnified=✗, #Params=7B, Visual Representation Type=Continuous2026.06 | 58.6 | — | — | — | — | — | |
| Gemini Pro 1.52024.09 | 58.5 | — | — | — | — | — | |
| Ovis2-8BSize=8B2025.12 | 57.4 | — | — | — | — | — | |
| Visual-SR12025.09 | 57.2 | — | — | — | — | — | |
| VisionZero-CLEVRBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 56.98 | — | — | — | — | — | |
| GPT-4V2024.01 | 56.8 | — | — | — | — | — | |
| GPT-4VLLM Size=Unk2024.03 | 56.8 | — | — | — | — | — | |
| GPT-4VAccess Type=Closed-source API2024.07 | 56.8 | — | — | — | — | — | |
| LLaVA-OneVisionSize=72B2024.09 | 56.8 | — | — | — | — | — | |
| GPT-4V2024.09 | 56.8 | — | — | — | — | — | |
| LLaVA-OneVision-72B2025.12 | 56.8 | — | — | — | — | — | |
| DeFacto (Ours)Backbone=Qwen2.5-VL-7B2025.09 | 56.6 | — | — | — | — | — | |
| Active-ZeroBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 56.42 | — | — | — | — | — | |
| Oryx-1.5Size=32B2024.09 | 56.1 | — | — | — | — | — | |
| Phi-4-Multimodal-5.6BSize=5.6B2025.12 | 56 | — | — | — | — | — | |
| InternVL2.5Unified=✗, #Params=8B, Visual Representation Type=Continuous2026.06 | 56 | — | — | — | — | — | |
| Qwen3-VL-2B-MastersSize=2B2025.12 | 55.8 | — | — | — | — | — | |
| LLaVA-OneVision-1.5-8BSize=8B2025.12 | 55.4 | — | — | — | — | — | |
| Qwen2.5-VLScale=32B2026.05 | 55.4 | — | — | — | — | — | |
| EvoLMMBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 55.31 | — | — | — | — | — | |
| BagelUnified=✓, #Params=7B, Visual Representation Type=Continuous2026.06 | 55.3 | — | — | — | — | — | |
| Qwen2.5-VL-7BSize=7B2025.12 | 55 | — | — | — | — | — | |
| Phi4-multimodal + DRScaffoldScale=5.6B2026.05 | 54.2 | — | — | — | — | — | |
| Qwen2-VL-7BParameters=7B2024.12 | 54.1 | — | — | — | — | — | |
| Molmo-72B2025.12 | 54.1 | — | — | — | — | — | |
| Phi4-multimodalScale=5.6B2026.05 | 54.1 | — | — | — | — | — | |
| GPT-4vModel Category=Proprietary VLMs2024.06 | 53.8 | — | — | — | — | — | |
| VotingAggregation Strategy=Majority Voting, Model Scale=7B-12B Ensemble2026.05 | 53.66 | — | — | — | — | — | |
| Base ModelBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 53.52 | — | — | — | — | — | |
| 360VLModel Size=70B, Access Type=Open-source2024.07 | 53.4 | — | — | — | — | — | |
| Qwen3-VL-2BSize=2B2025.12 | 53.4 | — | — | — | — | — | |
| Prism Captioner-7BModel Category=Prism Models, Reasoning Module=Llama32024.06 | 53.3 | — | — | — | — | — | |
| VisionZero-RealWorldBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 53.3 | — | — | — | — | — | |
| VisionZero-ChartBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 53.07 | — | — | — | — | — | |
| LLaVA-OneVision-1.5-4BSize=4B2025.12 | 52.7 | — | — | — | — | — | |
| VisPlayBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 52.51 | — | — | — | — | — | |
| Pixel Reasoner2025.09 | 52.5 | — | — | — | — | — | |
| Gemma-12BModel Backbone=Gemma, Model Scale=12B, Aggregation Strategy=Single Model2026.05 | 52.49 | — | — | — | — | — |