Visual Mathematical Reasoning on MathVista (score)
81.9ScoreGPT-5-Thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-ThinkingModel Category=Close-source Models2025.09 | 81.9 | |
| Gemini-2.5-ProModel Category=Close-source Models2025.09 | 80.9 | |
| CGC-8BBase MLLM=Qwen3-VL2026.04 | 78.2 | |
| VAPO-Thinker-7BModel Category=Our Models2025.09 | 75.6 | |
| Qwen3-VL-8B2026.04 | 75.3 | |
| Vision-R1-7BModel Category=Open-source Models2025.09 | 73.5 | |
| MM-Eureka-7BBase MLLM=Qwen2.5-VL2026.04 | 72 | |
| ThinkLite-VL-7BBase MLLM=Qwen2.5-VL2026.04 | 71.89 | |
| NoisyRollout-7BBase MLLM=Qwen2.5-VL2026.04 | 71.6 | |
| CGC-7BBase MLLM=Qwen2.5-VL2026.04 | 70.9 | |
| VLAA-Thinker-7BBase MLLM=Qwen2.5-VL2026.04 | 70.8 | |
| InternVL3-8BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 70.5 | |
| QvQ-72B-PreviewModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 70.3 | |
| CHARTOOL-7BParameters=7B2026.04 | 70.1 | |
| DRIFTModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reproduced by authors2025.10 | 69.9 | |
| X-REASONERModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 69 | |
| SFT BaselineModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reproduced by authors2025.10 | 68.7 | |
| InternVL2.5-8BModel Category=Open-source Models2025.09 | 68.2 | |
| Qwen2.5-VL-7BParameters=7B2026.04 | 68.2 | |
| VLAA-Thinker-7BModel Category=Open-source Models2025.09 | 68 | |
| HeadLensBackbone=Qwen3-VL-8B2026.03 | 67.9 | |
| MiCo-7BBase MLLM=Qwen2.5-VL2026.04 | 67.9 | |
| Qwen2.5-VL-7BModel Category=Open-source Models, Evaluation Source=Reproduced by authors2025.10 | 67.9 | |
| InternVL3-8B2026.04 | 67.3 | |
| VAPO-Thinker-3BModel Category=Our Models2025.09 | 67.1 | |
| Qwen2.5-VL-7B2026.04 | 67.1 | |
| BaselineBackbone=Qwen3-VL-8B2026.03 | 66.3 | |
| Kimi-VL-16BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 66 | |
| OpenVLThinker-7BModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 65.3 | |
| InternVL2.5-8BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 64.5 | |
| R1-OneVision-7BModel Category=Open-source Models2025.09 | 64.1 | |
| R1-Onevision-7BModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 64.1 | |
| InternLM-XComposer2.5Model Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 64 | |
| R1-VL-7BModel Category=Reasoning Fine-tuning Methods, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 63.5 | |
| LLaVA-OneVision-7BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 62.6 | |
| Qwen2.5-VL-7BModel Category=Open-source Models2025.09 | 62.3 | |
| Qwen2.5-VL-3BParameters=3B2026.04 | 62.3 | |
| Qwen2-VL-7BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 61.6 | |
| Qwen2-VL-7B2026.04 | 61.2 | |
| CHARTOOL-3BParameters=3B2026.04 | 60.5 | |
| Migician-7BBase MLLM=Qwen2-VL2026.04 | 60.2 | |
| HeadLensBackbone=Qwen2.5-VL-7B2026.03 | 58.3 | |
| InternVL2-8BModel Category=Open-source Models, Evaluation Source=Reported by Open Vision Reasoner (Wei et al., 2025)2025.10 | 58.3 | |
| BaselineBackbone=Qwen2.5-VL-7B2026.03 | 57.3 | |
| HeadLensBackbone=LLaVA-1.5-7B2026.03 | 26.3 | |
| BaselineBackbone=LLaVA-1.5-7B2026.03 | 23.9 |