Multimodal Reasoning on MMMU (val)
78.2AccuracyOpenAI-o1
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenAI-o1#Data=/2025.06 | 78.2 | |
| Claude-Sonnet-3.7Model Type=Proprietary2025.12 | 75 | |
| Qwen3-VL-8B-ThinkingModel Type=Reasoning2025.12 | 74.1 | |
| GPT-5-NanoModel Type=Proprietary2025.12 | 72.6 | |
| Keye-VL-1.5-8BModel Type=Reasoning2025.12 | 71.4 | |
| InternVL3.5-8B-MPOModel Type=Reasoning2025.12 | 71.2 | |
| Qwen3-VL-8B-Instruct + CAREBackbone=Qwen3-VL-8B-Instruct, Method=CARE2025.12 | 71 | |
| Claude-3.7-Sonnet#Data=/2025.06 | 71 | |
| Qwen3-VL-4B-ThinkingModel Type=Reasoning2025.12 | 70.8 | |
| Qwen2.5-VL-72BModel Category=General Multimodal LLM, Parameters=72B2025.12 | 70.2 | |
| Qwen2.5-VL-72B-IT#Data=/2025.06 | 70.2 | |
| InternVL2.5-78Bparameters=78B2024.12 | 70.1 | |
| Gemini-2.0-ProModel Type=Proprietary2025.12 | 69.9 | |
| Qwen3-VL-8B-InstructModel Type=Instruct2025.12 | 69.6 | |
| GPT-4oChain-of-Thought (CoT)=true2024.07 | 69.1 | |
| GPT-4o-202405132024.12 | 69.1 | |
| GPT-4o#Data=/2025.06 | 69.1 | |
| Claude 3.5Chain-of-Thought (CoT)=true2024.07 | 68.3 | |
| Claude-3.5-Sonnet2024.12 | 68.3 | |
| InternVL3.5-8B-InstructModel Type=Instruct2025.12 | 68.1 | |
| Qwen3-VL-4B-InstructModel Type=Instruct2025.12 | 67.4 | |
| Qwen3-VLModel Size=4B, Training Stage=instruct2026.05 | 67.4 | |
| MiMo-VL-7B-RLModel Type=Reasoning2025.12 | 66.7 | |
| InternVL-3.5Model Size=4B2026.05 | 66.6 | |
| Qwen3-VL-SegModel Size=4B, Training Stage=S-22026.05 | 66.2 | |
| MiMo-VL-7B-SFTModel Type=Instruct2025.12 | 64.6 | |
| Llama 3-V 405BChain-of-Thought (CoT)=true2024.07 | 64.5 | |
| Qwen2-VL-72Bparameters=72B2024.12 | 64.5 | |
| InternVL2.5-38Bparameters=38B2024.12 | 63.9 | |
| Qwen3-VLModel Size=4B, Training Stage=S-12026.05 | 63.4 | |
| GPT-4V2024.12 | 63.1 | |
| InternVL2-Llama3-76Bparameters=76B, backbone=Llama32024.12 | 62.7 | |
| Qwen2.5-VL-7B + CAREBackbone=Qwen2.5-VL-7B, Method=CARE2025.12 | 62.5 | |
| Gemini 1.5 ProChain-of-Thought (CoT)=true2024.07 | 62.2 | |
| Qwen2.5-VL-7B + GSPOBackbone=Qwen2.5-VL-7B, RL Method=GSPO2025.12 | 62.2 | |
| Gemini-1.5-Pro2024.12 | 62.2 | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B, RL Method=DAPO2025.12 | 61.6 | |
| Qwen2.5-VL-7BModel Type=Instruct2025.12 | 61.3 | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B, RL Method=GRPO2025.12 | 61.1 | |
| Perception-R1-7B#Data=1.4K2025.06 | 60.8 | |
| Llama 3-V 70BChain-of-Thought (CoT)=true2024.07 | 60.6 | |
| InternVL2.5-26Bparameters=26B2024.12 | 60 | |
| NVLM-D-72Bparameters=72B2024.12 | 59.7 | |
| Lumina-DiMOOModel Category=Unified Models2026.04 | 58.6 | |
| OpenVLThinker-7B#Data=25K2025.06 | 58.4 | |
| MM-Eureka-7B#Data=15K2025.06 | 58 | |
| ReLaX-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 57.4 | |
| InternVL3.5Parameters=4B2026.04 | 57.4 | |
| InternVL3.5Parameters=8B2026.04 | 57.2 | |
| SRPO-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 57.1 | |
| Intern-S1variant=mini, parameters=9B2026.01 | 57 | |
| InternVL3.5parameters=8B2026.01 | 56.89 | |
| LLaVA-OneVision-72Bparameters=72B2024.12 | 56.8 | |
| VL-Rethinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 56.7 | |
| SophiaVL-R1-7B#Data=130K2025.06 | 56.7 | |
| GPT-4VChain-of-Thought (CoT)=true2024.07 | 56.4 | |
| Intern2.5-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 56 | |
| InternVL2.5-8Bparameters=8B2024.12 | 56 | |
| InternVL2.5-8B#Data=/2025.06 | 56 | |
| Kimi-VL-16BModel Category=General Multimodal LLM, Parameters=16B2025.12 | 55.7 | |
| LLaVA-OVversion=1.5, parameters=8B2026.01 | 55.44 | |
| LLaVA-OneVision-1.5 8BModel Type=Instruct2025.12 | 55.4 | |
| BAGELModel Category=Unified Models2026.04 | 55.3 | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 55.22 | |
| MM-Eureka-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 55.2 | |
| InternVL2-40Bparameters=40B2024.12 | 55.2 | |
| Qwen2.5-VL-7B-IT#Data=/2025.06 | 55.2 | |
| Vision-R1-7B#Data=200K2025.06 | 55.2 | |
| VILA-1.5-40Bparameters=40B2024.12 | 55.1 | |
| Ovis1.6-Gemma2-9Bparameters=9B, backbone=Gemma22024.12 | 55 | |
| Vision-R1-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 54.7 | |
| VLAA-Thinker-7B#Data=25K2025.06 | 54.7 | |
| InternVL-UModel Category=Unified Models2026.04 | 54.7 | |
| BARD-VLParameters=8B, B=42026.04 | 54.6 | |
| Qwen2.5-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 54.3 | |
| Qwen2-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 54.1 | |
| Qwen2-VL-7Bparameters=7B2024.12 | 54.1 | |
| Molmo-72Bparameters=72B2024.12 | 54.1 | |
| Qwen3-VLParameters=8B2026.04 | 53 | |
| BARD-VLParameters=4B, B=322026.04 | 53 | |
| R1-OneVision-7B#Data=155K2025.06 | 52.9 | |
| Qwen2.5-VL-3B-InstructParameters=3B, Mode=Instruct2025.05 | 52.78 | |
| InternVL2-8Bparameters=8B2024.12 | 52.6 | |
| SFT-MModel=Qwen2.5-VL-7B, Paradigm=SFT-M2026.02 | 52.56 | |
| OpenVLThinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 52.5 | |
| InternVL2.5-4BComplexity=O(n)2025.12 | 52.3 | |
| InternVL2.5-4Bparameters=4B2024.12 | 52.3 | |
| R1-VL-7B#Data=260K2025.06 | 52.3 | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 52.22 | |
| ReLaX-VL-3BModel Category=Reasoning Multimodal LLM, Parameters=3B2025.12 | 52.2 | |
| GRPOModel=Qwen2.5-VL-7B, Paradigm=GRPO2026.02 | 51.89 | |
| SFTModel=Qwen2.5-VL-7B, Paradigm=SFT2026.02 | 51.67 | |
| Dream-VLParameters=7B2026.04 | 51.6 | |
| Qwen3-VLparameters=8B2026.01 | 51.44 | |
| Qwen2.5-VL-7BModel Category=Specialist VLMs2026.04 | 51.3 | |
| Intern2-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 51.2 | |
| InternVL2-26Bparameters=26B2024.12 | 51.2 | |
| SFT-RSModel=Qwen2.5-VL-7B, Paradigm=SFT-RS2026.02 | 50.33 | |
| SFT-MModel=Qwen2.5-VL-3B, Paradigm=SFT-M2026.02 | 50.11 | |
| LLaDA2.0-UniModel Category=Unified Models2026.04 | 50.1 |