Multimodal Mathematical Reasoning on MathVerse
62.4Average ScoreMathFlow-P-7B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=Gemini 2.5-pro2025.03 | 62.4 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=Claude-sonnet-3.52025.03 | 60.8 | — | — | — | — | — | |
| Gemini 2.5-proPerception Model=Gemini 2.5-pro2025.03 | 59.9 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=GPT-4o2025.03 | 59.5 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=GPT-4o-mini2025.03 | 58.9 | — | — | — | — | — | |
| GPT-4oPerception Model=GPT-4o2025.03 | 57.9 | — | — | — | — | — | |
| Claude-sonnet-3.5Perception Model=Claude-sonnet-3.52025.03 | 57.4 | — | — | — | — | — | |
| Vision-R1 (7B)Backbone=7B2025.08 | 57.3 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=InternVL-2.5-78B2025.03 | 56.8 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=GPT-4V2025.03 | 56.7 | — | — | — | — | — | |
| Vision-SR1Backbone=Qwen2.5-VL-7B2025.08 | 54.5 | — | — | — | — | — | |
| GPT-4VPerception Model=GPT-4V2025.03 | 54.4 | — | — | — | — | — | |
| Vision-R1Backbone=Qwen2.5-VL-7B, Training Data=47K data2025.08 | 53.2 | — | — | — | — | — | |
| GPT-4o-miniPerception Model=GPT-4o-mini2025.03 | 52.7 | — | — | — | — | — | |
| Percention-R1 (7B)Backbone=7B2025.08 | 52.1 | — | — | — | — | — | |
| Zero-shot InferenceBackbone=Qwen2.5-VL-7B, Inference Protocol=Zero-shot (before RL)2025.08 | 49.2 | — | — | — | — | — | |
| SophiaVL-R1-7BModel Category=Open-Source Reasoning MLLMs2025.05 | 48.8 | 45.4 | 43.9 | 45.1 | 58.5 | 51.3 | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=Qwen2-VL-72B2025.03 | 48.1 | — | — | — | — | — | |
| R1-OneVision-7BModel Category=Open-Source Reasoning MLLMs2025.05 | 46.4 | — | 40 | — | — | — | |
| Vision-SR1Backbone=Qwen2.5-VL-3B2025.08 | 45.8 | — | — | — | — | — | |
| URSA-8BModel Category=Open-Source Math MLLMs2025.05 | 45.7 | 46.4 | 34.6 | 43.9 | 55.3 | 48.3 | |
| VL-Rethinker-7BBackbone=Qwen2.5-VL-7B2026.05 | 45.68 | — | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPOModel Category=Open-Source Reasoning MLLMs, Training Strategy=GRPO2025.05 | 45.3 | 43 | 41 | 41.1 | 56 | 45.6 | |
| Visionary-R1 (3B)Backbone=3B2025.08 | 45 | — | — | — | — | — | |
| Qwen2.5-VL-7B + SRPOBackbone=Qwen2.5-VL-7B2026.05 | 44.92 | — | — | — | — | — | |
| Zero-shot InferenceBackbone=Qwen2.5-VL-3B, Inference Protocol=Zero-shot (before RL)2025.08 | 44.3 | — | — | — | — | — | |
| Qwen2.5-VL-7B-InstructModel Category=Open-Source Reasoning MLLMs2025.05 | 44 | 41.1 | 41 | 38.7 | 55.2 | 44 | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=Qwen-VL-MaX2025.03 | 43.3 | — | — | — | — | — | |
| InternVL-2.5-78BPerception Model=InternVL-2.5-78B2025.03 | 43.2 | — | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + SFT + GRPOModel Category=Open-Source Reasoning MLLMs, Training Strategy=SFT+GRPO2025.05 | 43.1 | 42.5 | 37.1 | 37.3 | 52.2 | 46.3 | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.05 | 43.02 | — | — | — | — | — | |
| Vision-R1Backbone=Qwen2.5-VL-3B, Training Data=47K data2025.08 | 42.8 | — | — | — | — | — | |
| Vision-Matters-7BBackbone=Qwen2.5-VL-7B2026.05 | 42.79 | — | — | — | — | — | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B2026.05 | 41.62 | — | — | — | — | — | |
| PAPO-G-7BBackbone=Qwen2.5-VL-7B2026.05 | 41.24 | — | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.05 | 40.86 | — | — | — | — | — | |
| Perception-R1-7BBackbone=Qwen2.5-VL-7B2026.05 | 40.35 | — | — | — | — | — | |
| Vision-SR1Backbone=Mimo-VL-7B2025.08 | 40 | — | — | — | — | — | |
| Qwen2.5-VL-3B w/ CADFTModel=Qwen2.5-VL-3B, Fine-tuning Protocol=CADFT2026.04 | 39.9 | 34.2 | 38.2 | 35.6 | — | — | |
| Qwen2.5-VL-3B + SRPOBackbone=Qwen2.5-VL-3B2026.05 | 39.72 | — | — | — | — | — | |
| Vision-SR1-7BBackbone=Qwen2.5-VL-7B2026.05 | 39.46 | — | — | — | — | — | |
| Qwen2-VL-72BPerception Model=Qwen2-VL-72B2025.03 | 38.9 | — | — | — | — | — | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B2026.05 | 38.71 | — | — | — | — | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.05 | 38.7 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=InfiMM-Math2025.03 | 38.1 | — | — | — | — | — | |
| Qwen2.5-VL-3B w/ DFTModel=Qwen2.5-VL-3B, Fine-tuning Protocol=DFT2026.04 | 37.54 | 32.49 | 35.91 | 33.5 | — | — | |
| ThinkLite-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 37.22 | — | — | — | — | — | |
| PAPO-D-3BBackbone=Qwen2.5-VL-3B2026.05 | 37.18 | — | — | — | — | — | |
| Qwen-VL-MaXPerception Model=Qwen-VL-MaX2025.03 | 36.2 | — | — | — | — | — | |
| Qwen2.5-VL-3B w/ SFTModel=Qwen2.5-VL-3B, Fine-tuning Protocol=SFT2026.04 | 35.66 | 30.96 | 33.63 | 32.74 | — | — | |
| PRCO-3BBackbone=Qwen2.5-VL-3B2026.05 | 35.53 | — | — | — | — | — | |
| Zero-shot InferenceBackbone=Mimo-VL-7B, Inference Protocol=Zero-shot (before RL)2025.08 | 35.5 | — | — | — | — | — | |
| Qwen2.5-VL-3B + DAPOBackbone=Qwen2.5-VL-3B2026.05 | 35.41 | — | — | — | — | — | |
| Vision-R1Backbone=Mimo-VL-7B, Training Data=47K data2025.08 | 35.3 | — | — | — | — | — | |
| InfiMM-MathPerception Model=InfiMM-Math2025.03 | 34.5 | — | — | — | — | — | |
| Vision-SR1-3BBackbone=Qwen2.5-VL-3B2026.05 | 34.01 | — | — | — | — | — | |
| PAPO-G-3BBackbone=Qwen2.5-VL-3B2026.05 | 33.88 | — | — | — | — | — | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL-3B, Fine-tuning Protocol=Base2026.04 | 33.83 | 28.81 | 30.96 | 31.6 | — | — | |
| Math-PUMA-Qwen2VL-7BModel Category=Open-Source Math MLLMs2025.05 | 33.6 | 33.4 | 26 | 31.6 | 42.1 | 35 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 33.37 | — | — | — | — | — | |
| Qwen2.5-VL-3B + GRPOBackbone=Qwen2.5-VL-3B2026.05 | 33.24 | — | — | — | — | — | |
| GPT-4VModel Category=General MLLMs2025.05 | 32.8 | — | — | — | — | — | |
| InternVL2.5-8B-VisualPRMModel Category=Open-Source Reasoning MLLMs2025.05 | 30.7 | 28.9 | 35.8 | 27.3 | 31.7 | 29.7 | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B2026.05 | 30.32 | — | — | — | — | — | |
| MathFlow-P-7BPerception Model=MathFlow-P-7B, Inference Model=InternLM-XC22025.03 | 30.2 | — | — | — | — | — | |
| LLaVA-OneVision-72BModel Category=General MLLMs2025.05 | 27.2 | — | — | — | — | — | |
| Multimath-7BModel Category=Open-Source Math MLLMs2025.05 | 26.9 | 28.1 | 15 | 25.9 | 34.8 | 30.8 | |
| LLaVA-OneVision-7BModel Category=General MLLMs2025.05 | 26.2 | — | — | — | — | — | |
| InternLM-XC2Perception Model=InternLM-XC22025.03 | 25.9 | — | — | — | — | — | |
| Math-LLaVA-13BModel Category=Open-Source Math MLLMs2025.05 | 22.9 | 24.5 | 16.1 | 21.7 | 27.3 | 24.9 |