Mathematical Multimodal Reasoning on MathVista (Accuracy)
85.6AccuracySeed-1.5-thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Seed-1.5-thinkingAccess Type=Closed-Source2025.12 | 85.6 | |
| Gemini-2.5-Pro-ThinkingAccess Type=Closed-Source2025.12 | 83.8 | |
| Gemini 2.5 ProTool=Code, Param Size=-2025.11 | 83 | |
| MM-Eureka-32B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 80.2 | |
| InternVL3.5-8B-MPOTraining=DF-GRPO2026.03 | 79.9 | |
| Qwen2.5-VL-7B-Instruct + R-TAPModel Size=7B2026.03 | 79.4 | |
| MM-Eureka-7B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 79.3 | |
| InternVL3-78BModel Category=Open-Source2025.06 | 79 | |
| Qwen2.5-VL-32BTraining=DF-GSPO2026.03 | 78.8 | |
| InternVL3.5-8B-MPOTraining=PRIME2026.03 | 78.5 | |
| InternVL3.5-8B-MPOTraining=GRPO2026.03 | 78.4 | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 78.2 | |
| R1-ShareVL-32B2026.03 | 77.6 | |
| InternVL-3.5-4BModel Scale=4B2026.02 | 77.1 | |
| Qwen2.5-VL-7B (Bo8)Training Strategy=GRPO, Sampling Strategy=Best-of-82025.06 | 76.8 | |
| Qwen2.5-VL-32BReasoning Framework=Standard2025.06 | 76.8 | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 76.8 | |
| InternVL2.5-78B-MPOModel Category=Open-source Models (>70B)2025.06 | 76.6 | |
| DenseMLLM-4BModel Scale=4B2026.02 | 76.5 | |
| Vision-G1Access Type=Open-Source, Model Scale=7B2025.12 | 76.1 | |
| Ovis2-34BModel Category=Open-Source2025.06 | 76.1 | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 76.1 | |
| Vision-G12026.03 | 76.1 | |
| Qwen2.5-VL-7BTraining=DF-GSPO2026.03 | 76.1 | |
| PeBR-R1Access Type=Open-Source, Model Scale=7B2025.12 | 76 | |
| ETTCAggregation Strategy=ETTC, Model Scale=7B-12B Ensemble2026.05 | 75.93 | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 75.9 | |
| InternVL3.5-8B-MPOTraining=MPO2026.03 | 75.9 | |
| SRPO-7BMethod Category=Verification-augmented2025.06 | 75.8 | |
| Sora-2 AudioInput Modality=Audio, LLM-as-a-Judge=GPT-4o2025.11 | 75.7 | |
| VAPO-ThinkerAccess Type=Open-Source, Model Scale=7B2025.12 | 75.6 | |
| Qwen2.5-VL-32BTraining=GSPO2026.03 | 75.5 | |
| R1-ShareVL-7BEvaluation Source=original paper2026.03 | 75.4 | |
| R1-ShareVL-7B2026.03 | 75.4 | |
| ThinkLite-VLAccess Type=Open-Source, Model Scale=7B2025.12 | 75.1 | |
| InternVL3-14BModel Category=Open-Source2025.06 | 75.1 | |
| InternVL3-38BModel Category=Open-Source2025.06 | 75.1 | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 75.1 | |
| VL-RethinkerAccess Type=Open-Source, Model Scale=7B2025.12 | 74.9 | |
| VL-Rethinker-7BModel Category=Open-Source2025.06 | 74.9 | |
| Qwen2.5-VL-72BModel Category=Larger MLLMs without Reasoning2025.12 | 74.8 | |
| Qwen-2.5-VL-72B2026.03 | 74.8 | |
| Qwen-2.5-VL-72BModel Category=Open-Source General Models2026.03 | 74.8 | |
| MM-Eureka-32BModel Category=Open-Source Reasoning Models2026.03 | 74.8 | |
| Qwen-2.5-VL-32BEvaluation Source=original paper2026.03 | 74.7 | |
| Qwen-2.5-VL-32BModel Category=Open-Source General Models2026.03 | 74.7 | |
| MM-Eureka-32BModel Category=Open-Source2025.06 | 74.7 | |
| Qwen2.5-VL-32BTraining=Base2026.03 | 74.7 | |
| SPOT-EBase Model=InternVL3-8B, Source Type=Open-Source, Test-time Adaptation=SPOT-E2026.06 | 74.4 | |
| ThinkLite-7B2026.03 | 74.3 | |
| DIVA-GRPO-7B2026.03 | 74.2 | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 74.2 | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 74.2 | |
| ToR-DAPOData Size=39K (ViRL-39K), Base Model=Qwen-2.5-VL-7B2026.03 | 74.2 | |
| DR-MMSearchAgent (w/ SPAI)Model Category=Open-source Models2026.04 | 74.2 | |
| Qwen2.5-VL-72BModel Category=Open-source Models (>70B)2025.06 | 74.2 | |
| o12026.03 | 73.9 | |
| o1Model Category=Closed-Source Models2026.03 | 73.9 | |
| InternVL2.5-38B-MPO2026.03 | 73.8 | |
| InternVL2.5-38B-MPOModel Category=Open-Source Reasoning Models2026.03 | 73.8 | |
| Qwen3-VL-4BModel Scale=4B2026.02 | 73.7 | |
| VL-RethinkerTool=✗, Param Size=7B2025.11 | 73.7 | |
| Ovis2-16BModel Category=Open-Source2025.06 | 73.7 | |
| VL-Rethinker-7BModel Category=Open-source Models2026.04 | 73.7 | |
| InternVL3-8BModel Category=Open-Source2025.06 | 73.6 | |
| Qwen2.5-VL-7BTraining=GSPO2026.03 | 73.6 | |
| VisonR1 7BModel Category=Reasoning MLLMs2026.01 | 73.5 | |
| Adora-7B2026.03 | 73.5 | |
| R1-ShareVL-7BEvaluation Source=re-evaluation2026.03 | 73.5 | |
| ADORA-7BModel Category=Open-Source Reasoning Models2026.03 | 73.5 | |
| Vision-R1-7BModel Size=7B2026.03 | 73.5 | |
| Vision-R1-7BData Size=200K+10K, Base Model=Qwen-2.5-VL-7B2026.03 | 73.5 | |
| Vision-R12026.03 | 73.5 | |
| OursAccess Type=Open-Source, Model Scale=7B2025.12 | 73.4 | |
| MiniCPM-o-2.6-8BModel Category=Open-source Models (~7B)2025.06 | 73.3 | |
| SPOT-EBase Model=Qwen3-VL-8B, Source Type=Open-Source, Test-time Adaptation=SPOT-E2026.06 | 73.3 | |
| Revisual-R1Access Type=Open-Source, Model Scale=7B2025.12 | 73.1 | |
| MM-EurekaAccess Type=Open-Source, Model Scale=7B2025.12 | 73 | |
| MathBookAccess Type=Open-Source, Model Scale=7B2025.12 | 73 | |
| MM-Eureka-7BEvaluation Source=original paper2026.03 | 73 | |
| MM-Eureka-7BModel Category=Open-Source Reasoning Models2026.03 | 73 | |
| MM-Eureka-7BModel Category=Open-Source2025.06 | 73 | |
| Qwen2.5-VL-7B-Instruct + NoisyRolloutModel Size=7B2026.03 | 72.9 | |
| ThinkLite-7B-VLModel Size=7B2026.03 | 72.7 | |
| ThinkLite-7B-VLData Size=11K, Base Model=Qwen-2.5-VL-7B2026.03 | 72.7 | |
| SFTTraining Stage=SFT, RL Alignment (GRPO)=true2025.08 | 72.7 | |
| MM-EurekaTool=✗, Param Size=7B2025.11 | 72.6 | |
| NoisyRollout-7BData Size=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 72.6 | |
| ToR-DAPOData Size=2.1K (Geo3K), Base Model=Qwen-2.5-VL-7B2026.03 | 72.6 | |
| Claude Sonnet 4.5Input Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 72.5 | |
| ThinkLite-VL-7BModel Category=Open-source Models2026.04 | 72.4 | |
| OpenVLThinkerAccess Type=Open-Source, Model Scale=7B2025.12 | 72.3 | |
| Look-BackAccess Type=Open-Source, Model Scale=7B2025.12 | 72.3 | |
| InternVL2.5-VL-78B2026.03 | 72.3 | |
| InternVL2.5-VL-78BModel Category=Open-Source General Models2026.03 | 72.3 | |
| PSFTTraining Stage=PSFT, RL Alignment (GRPO)=true2025.08 | 72.2 | |
| Qwen-7BModel Backbone=Qwen, Model Scale=7B, Aggregation Strategy=Single Model2026.05 | 72.08 | |
| GPT-4.1Access Type=Closed-Source2025.12 | 72 | |
| InternVL2.5-VL-38B2026.03 | 71.9 | |
| InternVL2.5-VL-38BModel Category=Open-Source General Models2026.03 | 71.9 |