Multimodal Mathematical Reasoning on MathVista MINI
0.878AccuracyQwen3.5-27B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-27BMode=REASONING, Architecture=Dense, # Total Params=27B, # Activated Params=27B2026.04 | 0.878 | |
| Qwen3-VL-32BMode=Thinking, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 0.859 | |
| Qwen3-VL-235B-A22BMode=Thinking, Architecture=MoE, # Total Params=236B, # Activated Params=23B2026.04 | 0.858 | |
| EXAONE 4.5 33BMode=REASONING, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 0.85 | |
| Qwen3-VL-8B-Instruct + CAREBackbone=Qwen3-VL-8B-Instruct, Method=CARE2025.12 | 0.821 | |
| MiMo-VL-7B-SFTModel Type=Instruct2025.12 | 0.818 | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 0.818 | |
| MiMo-VL-7B-RLModel Type=Reasoning2025.12 | 0.815 | |
| Qwen3-VL-8B-ThinkingModel Type=Reasoning2025.12 | 0.814 | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 0.808 | |
| Keye-VL-1.5-8BModel Type=Reasoning2025.12 | 0.807 | |
| Qwen3-VL-4B-ThinkingModel Type=Reasoning2025.12 | 0.795 | |
| GPT-5 miniMode=REASONING: HIGH, Architecture=-, # Total Params=-, # Activated Params=-2026.04 | 0.791 | |
| InternVL3.5-8BModel Scale=8B, Optimization Method=N/A2026.06 | 0.784 | |
| MetisModel Category=Agentic Multimodal Models2026.04 | 0.78 | |
| Qwen3-VL-8B-InstructModel Type=Instruct2025.12 | 0.772 | |
| InternVL-3.5Model Size=4B2026.05 | 0.771 | |
| Qwen3-VL-8B-InstructModel Category=Open-source Models2026.04 | 0.763 | |
| Qwen2.5-VL-72B-IT#Data=/2025.06 | 0.758 | |
| Qwen3-VL-SegModel Size=4B, Training Stage=S-22026.05 | 0.755 | |
| ThinkLite-VL-7BModel Category=Text-only Reasoning Models2026.04 | 0.751 | |
| M2-ReasoningModel Type=Reasoning2025.12 | 0.75 | |
| VL-Rethinker-7BModel Category=Text-only Reasoning Models2026.04 | 0.749 | |
| Qwen2.5-VL-7B + CAREBackbone=Qwen2.5-VL-7B, Method=CARE2025.12 | 0.747 | |
| InternVL3.5-8B-InstructModel Type=Instruct2025.12 | 0.742 | |
| InternVL3.5-8B-MPOModel Type=Reasoning2025.12 | 0.742 | |
| Perception-R1-7B#Data=1.4K2025.06 | 0.742 | |
| Qwen2.5-VL-7B + GSPOBackbone=Qwen2.5-VL-7B, RL Method=GSPO2025.12 | 0.741 | |
| OpenAI-o1#Data=/2025.06 | 0.739 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B2026.01 | 0.737 | |
| Qwen3-VL-4B-InstructModel Type=Instruct2025.12 | 0.737 | |
| Qwen3-VLModel Size=4B, Training Stage=instruct2026.05 | 0.737 | |
| Vision-R1-7BModel Type=Reasoning2025.12 | 0.735 | |
| InternVL3.5-4B (+OPD)Model Scale=4B, Optimization Method=OPD2026.06 | 0.733 | |
| GPT-5-NanoModel Type=Proprietary2025.12 | 0.731 | |
| Vision-R1-7B#Data=200K2025.06 | 0.731 | |
| InternVL3.5-4B (+GNDPO)Model Scale=4B, Optimization Method=GNDPO2026.06 | 0.73 | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B, RL Method=DAPO2025.12 | 0.726 | |
| MM-Eureka-7BModel Category=Text-only Reasoning Models2026.04 | 0.726 | |
| MM-Eureka-7B#Data=15K2025.06 | 0.725 | |
| D2IparReasoning Strategy=D2I, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 0.722 | |
| DeepEyesV2Model Category=Agentic Multimodal Models2026.04 | 0.719 | |
| VLAA-Thinker-7BModel Category=Text-only Reasoning Models2026.04 | 0.717 | |
| InternVL3-8BModel Category=Open-source Models2026.04 | 0.716 | |
| InternVL3.5-4B (Base (Instruct))Model Scale=4B, Optimization Method=Base (Instruct)2026.06 | 0.714 | |
| InternVL3.5-4B (+GSPO)Model Scale=4B, Optimization Method=GSPO2026.06 | 0.714 | |
| Gemini-2.0-ProModel Type=Proprietary2025.12 | 0.713 | |
| OpenVLThinker-7B#Data=25K2025.06 | 0.713 | |
| Qwen3-VLModel Size=4B, Training Stage=S-12026.05 | 0.709 | |
| VLAA-Thinker-7B#Data=25K2025.06 | 0.707 | |
| SophiaVL-R1-7B#Data=130K2025.06 | 0.706 | |
| OpenVLThinker-7B2025.07 | 0.702 | |
| DeepEyesModel Type=Reasoning2025.12 | 0.701 | |
| DeepEyesModel Category=Agentic Multimodal Models2026.04 | 0.701 | |
| ThymeModel Category=Agentic Multimodal Models2026.04 | 0.7 | |
| D2IlocReasoning Strategy=D2I, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 0.699 | |
| D2IjusReasoning Strategy=D2I, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 0.697 | |
| LLaVA-OneVision-1.5 8BModel Type=Instruct2025.12 | 0.696 | |
| Qwen2.5-VL-7B +PCBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 0.696 | |
| Qwen2.5-VL-7B w/ GRPO†Reasoning Mode=Intuitive, Backbone=Qwen2.5-VL-7B2025.07 | 0.695 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.01 | 0.694 | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B, RL Method=GRPO2025.12 | 0.689 | |
| Qwen2.5-VL-7BModel Type=Instruct2025.12 | 0.686 | |
| Qwen-2.5-VL-7B-InstructModel Category=Open-source Models2026.04 | 0.683 | |
| Qwen2.5-VL-7BBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=No2026.03 | 0.682 | |
| Qwen2.5-VL-7B*Backbone=Qwen2.5-VL-7B2025.07 | 0.682 | |
| Qwen2.5-VL-7B-IT#Data=/2025.06 | 0.681 | |
| Qwen2.5-VL-7B w/ GRPOReasoning Mode=Deliberate, Backbone=Qwen2.5-VL-7B2025.07 | 0.681 | |
| D2DlocReasoning Strategy=D2D, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 0.68 | |
| FineViT-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 0.677 | |
| Claude-Sonnet-3.7Model Type=Proprietary2025.12 | 0.668 | |
| Claude-3.7-Sonnet#Data=/2025.06 | 0.668 | |
| Qwen2.5-VL-3B + CAREBackbone=Qwen2.5-VL-3B, Method=CARE2025.12 | 0.665 | |
| InternVL3.5-2B (+GNDPO)Model Scale=2B, Optimization Method=GNDPO2026.06 | 0.66 | |
| D2DjusReasoning Strategy=D2D, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 0.657 | |
| R1-OneVision-7B#Data=155K2025.06 | 0.65 | |
| InternVL3.5-2B (+OPD)Model Scale=2B, Optimization Method=OPD2026.06 | 0.647 | |
| InternVL2.5-8B#Data=/2025.06 | 0.644 | |
| InternVL2.5-8B2025.07 | 0.644 | |
| R1-Onevision-7B2025.07 | 0.641 | |
| GPT-4oModel Type=Proprietary2025.12 | 0.638 | |
| GPT-4o#Data=/2025.06 | 0.638 | |
| GPT-4o2025.07 | 0.638 | |
| Qwen2.5-VL-3B +PCBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=Yes2026.03 | 0.633 | |
| R1-VL-7B#Data=260K2025.06 | 0.627 | |
| D2DparReasoning Strategy=D2D, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 0.626 | |
| Qwen2.5-VL-3BBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=No2026.03 | 0.623 | |
| InternVL3.5-2B (+GSPO)Model Scale=2B, Optimization Method=GSPO2026.06 | 0.623 | |
| Qwen2.5-VL-3B-InstructParams=3.75B2025.12 | 0.62 | |
| Qwen2.5-VL-3BModel Type=Instruct2025.12 | 0.62 | |
| Intern3.5-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 0.615 | |
| InternVL3.5-2B (Base (Instruct))Model Scale=2B, Optimization Method=Base (Instruct)2026.06 | 0.608 | |
| MobileNet-QwenParams=1.84B2025.12 | 0.606 | |
| URSA-7B#Data=3.06M2025.06 | 0.598 | |
| Aquila-VLLanguage Backbone=Qwen2.5 1.5B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 0.593 | |
| Qwen2-VL-7B-IT#Data=/2025.06 | 0.586 | |
| LLaVA-OneVisionModel Category=Open-source Models2026.04 | 0.586 | |
| InternVL2-8B2025.07 | 0.583 | |
| Qwen2-VL-7B2025.07 | 0.582 | |
| GPT-4V2025.07 | 0.581 |