Multimodal Reasoning on MMMU
83.89AccuracyGemini-2.5 (Pro)
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5 (Pro)Model Tier=Pro2026.01 | 83.89 | |
| S1-VL-32B-RLParameter Count=32B, Training Phase=RL2026.04 | 83.4 | |
| S1-VL-32B-SFTParameter Count=32B, Training Phase=SFT2026.04 | 82.5 | |
| Claude Sonnet 4.5Input Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 82 | |
| Gemini 2.5 Pro2026.04 | 82 | |
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 81.4 | |
| GPT-52026.04 | 81.22 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 80.55 | |
| STEP3-VL-10B (PaCoRe)Reasoning Strategy=PaCoRe, Parameters=10B2026.01 | 80.11 | |
| Seed-1.5-VL (Thinking)Thinking Mode=true2026.01 | 79.11 | |
| Gemini 2.5 ProInput Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 79 | |
| Qwen3-VL (Thinking)Thinking Mode=true, Parameters=235B-A22B2026.01 | 78.7 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 78.4 | |
| STEP3-VL-10BNumber of Parameters=10B2026.01 | 78.11 | |
| STEP3-VL-10B (SeRe)Reasoning Strategy=SeRe, Parameters=10B2026.01 | 78.11 | |
| Qwen3-VL-235B-A22B-ThinkingParameter Count=235B-A22B, Reasoning Mode=Thinking2026.04 | 77.89 | |
| GPT-5 highInput Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 77 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 76.56 | |
| Qwen3-VL-32B-ThinkingParameter Count=32B, Reasoning Mode=Thinking2026.04 | 76 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 76 | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 75.8 | |
| Intern-S1Parameter Count=235B+6B2026.04 | 75.56 | |
| GLM-4.6VParameters=106B-A12B2026.01 | 75.2 | |
| Claude-3.7-SonnetModel Category=Proprietary2025.06 | 75 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 74.78 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 74.1 | |
| Qwen3-VL ThinkingNumber of Parameters=8B2026.01 | 73.53 | |
| Gemini 2.5 Flash2026.04 | 72.7 | |
| Gemini-2.0-FlashModel Category=Proprietary2025.06 | 72.6 | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 72.4 | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 72.3 | |
| InternVL3-78BModel Category=Open-Source2025.06 | 72.2 | |
| InternVL 3.5Number of Parameters=8B2026.01 | 71.69 | |
| GLM-4.6V FlashNumber of Parameters=9B2026.01 | 71.17 | |
| MiMo-VL RL-2508Number of Parameters=7B2026.01 | 71.14 | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 71.1 | |
| GPT-4o-20241120Model Category=Proprietary2025.06 | 70.7 | |
| Gemini-2.0-FlashActivation Replay=false2025.11 | 70.7 | |
| LongCat-NextParameters=68BA3B, Internal Reasoning (Think mode)=false2026.05 | 70.6 | |
| QVQ-72B-PreviewModel Category=Open-Source2025.06 | 70.3 | |
| QvQ-72B-PreviewActivation Replay=false2025.11 | 70.3 | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 70.2 | |
| InternVL3-38BModel Category=Open-Source2025.06 | 70.1 | |
| Intern-S1-miniParameter Count=8B2026.04 | 70 | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 69.8 | |
| ICLAModel=Qwen2.5-VL-7B2026.02 | 69.2 | |
| Sora-2 AudioInput Modality=Audio, LLM-as-a-Judge=GPT-4o2025.11 | 69.2 | |
| Qwen2.5-VL-32BReasoning Framework=Standard2025.06 | 69 | |
| VCDModel=Qwen2.5-VL-7B2026.02 | 68.3 | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 68.2 | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 68.2 | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 67.8 | |
| VanillaModel=Qwen2.5-VL-7B2026.02 | 67.5 | |
| RLSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 67.22 | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 67.2 | |
| InternVL3-14BModel Category=Open-Source2025.06 | 67.1 | |
| Ovis2-34BModel Category=Open-Source2025.06 | 66.7 | |
| VDDModel=Qwen2.5-VL-7B2026.02 | 65.8 | |
| DAMOModel=Qwen2.5-VL-7B2026.02 | 65.8 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 65.11 | |
| SDPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 65.11 | |
| Gemma-3-27BParameters=27B2025.12 | 64.9 | |
| Gemma-3-27BModel Category=Open-Source2025.06 | 64.9 | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 64.7 | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 64.6 | |
| OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 63.82 | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 63.8 | |
| GRPO + OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 63.22 | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 63.2 | |
| InternVL3-8BParameters=8B2025.12 | 62.7 | |
| InternVL3-8BModel Category=Open-Source2025.06 | 62.7 | |
| InternVL3-8BActivation Replay=false2025.11 | 62.7 | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 62.7 | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 62.6 | |
| DeCoModel=Qwen2.5-VL-7B2026.02 | 62.5 | |
| Base LLMBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 62.44 | |
| OmniCaptionerMethod Category=Caption-then-Reason2025.06 | 62.2 | |
| MM-Eureka-32BModel Category=Open-Source2025.06 | 62 | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 62 | |
| ECSOMethod Category=Caption-then-Reason2025.06 | 61.4 | |
| Qwen2.5-VL-3B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 60.9 | |
| DoLAModel=Qwen2.5-VL-7B2026.02 | 60.8 | |
| Ovis2-16BModel Category=Open-Source2025.06 | 60.7 | |
| Athena-PRMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 60.3 | |
| VisualPRM-8BPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 60.2 | |
| Self-consistencyPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 60.1 | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 60 | |
| Gemini Ultra2023.02 | 59.4 | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 59.3 | |
| Athena-ORMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 59.1 | |
| InternVL3.5-2B2025.12 | 59 | |
| InternVL3.5Number of Parameters=2.3B2025.12 | 59 | |
| InternVL3.5#Params=2B2026.03 | 59 | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 58.7 | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 58.7 | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 58.7 | |
| SPARK-VL-7BEvaluation Source=original paper2026.05 | 58.7 | |
| Qwen2.5-VL-7BActivation Replay=false2025.11 | 58.6 | |
| Qwen2.5Type=AR2026.05 | 58.6 | |
| Qwen2.5-VL-7B-InstructEvaluation Source=original paper2026.05 | 58.6 |