Multimodal Reasoning on MMStar
82AccuracyMasters
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MastersBase Model=InternVL3-8B2025.12 | 82 | — | |
| MastersBase Model=InternVL3.5-8B2025.12 | 80.8 | — | |
| MastersBase Model=Qwen3-VL-8B2025.12 | 79.7 | — | |
| Qwen3-VL-32B2025.12 | 77.7 | — | |
| GPT-52025.12 | 75.7 | — | |
| GLM-4.5V2025.12 | 75.3 | — | |
| InternVL3.5-38B2025.12 | 75.3 | — | |
| MastersBase Model=Qwen2.5-VL-7B2025.12 | 74.9 | — | |
| Qwen3-VL-32B-InstructContext Window=32K2026.03 | 74.3 | — | |
| GPT-5-Mini2025.12 | 74.1 | — | |
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 74.07 | — | |
| Qwen3-VL-32B-InstructContext Window=4K2026.03 | 73.7 | — | |
| Gemini-2.5-Pro2025.12 | 73.6 | — | |
| InternVL3-78B2025.12 | 72.5 | — | |
| Qwen2.5-VL-32BTraining=DF-GSPO2026.03 | 71.5 | — | |
| Qwen2.5-VL-72B2025.12 | 70.8 | — | |
| Qwen2.5-VL-72BModel Category=General Multimodal LLM, Parameters=72B2025.12 | 70.8 | — | |
| R1-ShareVL-32B2026.03 | 70.2 | — | |
| Qwen3-VL-8B-InstructContext Window=32K2026.03 | 69.9 | — | |
| GPT-4.12025.12 | 69.8 | — | |
| Qwen2.5-VL-32BTraining=Base2026.03 | 69.5 | — | |
| Claude-4-Sonnet2025.12 | 69.4 | — | |
| Gemini-2.0-Flash2025.12 | 69.4 | — | |
| Qwen3-VL-8B-InstructContext Window=4K2026.03 | 68.9 | — | |
| Qwen2.5-VL-7BTraining=DF-GSPO2026.03 | 68.3 | — | |
| Qwen2.5-VL-32BTraining=GSPO2026.03 | 68.3 | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 67.8 | — | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.03 | 67.2 | — | |
| R1-ShareVL-7B2026.03 | 67 | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.03 | 66.93 | — | |
| Qwen2.5vl-InstructModel Scale=7B, Training Protocol=PSO, Category=Open-Source Reasoning MLLMs2025.12 | 66.5 | — | |
| LLaVA-OneVisionModel Scale=72B, Category=Open-Source General MLLMs2025.12 | 66.1 | — | |
| Vision-G12026.03 | 66 | — | |
| Qwen2.5-VL-7BTraining=GSPO2026.03 | 65.9 | — | |
| LLaVA-OneVision-72B2025.12 | 65.8 | — | |
| ReLaX-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 65.5 | — | |
| DAPOBackbone=Qwen2.5-VL-7B2026.03 | 65.4 | — | |
| Vision-SR1-7BBackbone=Qwen2.5-VL-7B2026.03 | 65.26 | — | |
| Claude-3.5-Sonnet2025.12 | 65.1 | — | |
| Claude-3.7-Sonnet2025.12 | 65.1 | — | |
| R1-ShareVL-7BBackbone=Qwen2.5-VL-7B2026.03 | 65.06 | — | |
| PAPO-G-7BBackbone=Qwen2.5-VL-7B2026.03 | 64.93 | — | |
| Qwen2.5vl-InstructModel Scale=7B, Training Protocol=SFT + GRPO, Category=Open-Source Reasoning MLLMs2025.12 | 64.8 | — | |
| MMR1-7B-RLBackbone=Qwen2.5-VL-7B2026.03 | 64.8 | — | |
| GPT-4o2025.12 | 64.7 | — | |
| Phi-4-reasoning-vision-15B2026.03 | 64.5 | — | |
| Perception-R1-7BBackbone=Qwen2.5-VL-7B2026.03 | 64.33 | — | |
| Qwen2.5vl-InstructModel Scale=7B, Category=Open-Source Reasoning MLLMs2025.12 | 64.3 | — | |
| MM-Eureka-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 64.3 | — | |
| Base ModelBackbone=Qwen2.5-VL-7B2026.03 | 64.26 | — | |
| VL-Rethinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 64.2 | — | |
| GRPOBackbone=Qwen2.5-VL-7B2026.03 | 64.06 | — | |
| Active-ZeroBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 64 | — | |
| NoisyRollout-7BBackbone=Qwen2.5-VL-7B2026.03 | 63.93 | — | |
| Qwen2.5-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 63.9 | — | |
| Phi-4-reasoning-vision-15BProtocol=force nothink2026.03 | 63.9 | — | |
| Qwen2.5-VL-7BTraining=Base2026.03 | 63.9 | — | |
| NVLM-72B2025.12 | 63.7 | — | |
| ThinkLite-7B2026.03 | 63.7 | — | |
| InternVL2.5-8B-VisualPRMModel Scale=8B, Category=Open-Source Reasoning MLLMs2025.12 | 63.4 | — | |
| Molmo-72B2025.12 | 63.3 | — | |
| VisPlayBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 62.93 | — | |
| AStar (Qwen2.5-7B)OS Only=✓, Training-Free=✓, Prior Data=0.5K, Pre. Time=50 mins2025.02 | 62.3 | — | |
| VisionZero-ChartBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 62.27 | — | |
| VisionZero-RealWorldBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 62.27 | — | |
| Vision-Matters-7BBackbone=Qwen2.5-VL-7B2026.03 | 62.2 | — | |
| VisionZero-CLEVRBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 62.07 | — | |
| AStar (Qwen2-VL-7B)OS Only=✓, Training-Free=✓, Prior Data=0.5K, Pre. Time=50 mins2025.02 | 62 | — | |
| LLaVA-OneVisionModel Scale=7B, Category=Open-Source General MLLMs2025.12 | 61.7 | — | |
| Base ModelBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 61.53 | — | |
| EvoLMMBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 61.53 | — | |
| Intern2-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 61.5 | — | |
| Mulberry-7BType=Search, OS Only=✗, Training-Free=✗, Prior Data=260K2025.02 | 61.3 | — | |
| PRCO-3BBackbone=Qwen2.5-VL-3B2026.03 | 61 | — | |
| Vision-R1-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 60.9 | — | |
| Qwen2-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 60.7 | — | |
| ReLaX-VL-3BModel Category=Reasoning Multimodal LLM, Parameters=3B2025.12 | 60.7 | — | |
| PAPO-D-3BBackbone=Qwen2.5-VL-3B2026.03 | 60.66 | — | |
| DAPOBackbone=Qwen2.5-VL-3B2026.03 | 60.4 | — | |
| R1-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 60 | — | |
| R1-VL-7BType=GRPO, OS Only=✓, Training-Free=✗, Prior Data=260K2025.02 | 60 | — | |
| Kimi-VL-A3B-Instruct2026.03 | 60 | — | |
| MM-Eureka-7BType=GRPO, OS Only=✓, Training-Free=✗, Prior Data=15K2025.02 | 59.4 | — | |
| gemma-3-12b-it2026.03 | 59.4 | — | |
| PixelReasoner-7B + OursBackbone=PixelReasoner-7B, Modulated thinking process=True2026.03 | 59.4 | — | |
| R1-OnevisionModel Scale=7B, Category=Open-Source Reasoning MLLMs2025.12 | 59.1 | — | |
| Gemini-1.5-Pro2025.12 | 59.1 | — | |
| PAPO-G-3BBackbone=Qwen2.5-VL-3B2026.03 | 58.86 | — | |
| PixelReasoner-7BBackbone=PixelReasoner-7B2026.03 | 58.7 | — | |
| LMM-R1-3BType=PPO, OS Only=✓, Training-Free=✗, Prior Data=55.3K2025.02 | 58 | — | |
| GRPOBackbone=Qwen2.5-VL-3B2026.03 | 58 | — | |
| MMR1-3B-RLBackbone=Qwen2.5-VL-3B2026.03 | 57.2 | — | |
| GPT-4VCategory=Open-Source General MLLMs2025.12 | 57.1 | — | |
| Qwen2.5-VL-3B-InstructParams=3.75B2025.12 | 56.73 | — | |
| Vision-SR1-3BBackbone=Qwen2.5-VL-3B2026.03 | 56.73 | — | |
| OpenVLThinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 56.3 | — | |
| Base ModelBackbone=Qwen2.5-VL-3B2026.03 | 56.06 | — | |
| Qwen2.5-VL-3BModel Category=General Multimodal LLM, Parameters=3B2025.12 | 55.9 | — | |
| URSA-8BType=SFT, OS Only=✗, Training-Free=✗, Prior Data=1100K, Pre. Time=3 days2025.02 | 55.4 | — | |
| MobileNet-QwenParams=1.84B2025.12 | 55.07 | — |