Mathematical Reasoning on MathVision
75.95AccuracySTEP3-VL-10B (PaCoRe)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| STEP3-VL-10B (PaCoRe)Reasoning Strategy=PaCoRe, Parameters=10B2026.01 | 75.95 | — | — | — | — | — | |
| Gemini-2.5 (Pro)Model Tier=Pro2026.01 | 73.3 | — | — | — | — | — | |
| Qwen3-VL (Thinking)Thinking Mode=true, Parameters=235B-A22B2026.01 | 72.1 | — | — | — | — | — | |
| GPT-5-Thinking*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 72 | — | — | — | — | — | |
| STEP3-VL-10B (SeRe)Reasoning Strategy=SeRe, Parameters=10B2026.01 | 70.81 | — | — | — | — | — | |
| Gemini-2.5-Pro*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 69.1 | — | — | — | — | — | |
| Human2025.02 | 68.8 | — | — | — | 61.3 | — | |
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 68.75 | — | — | — | — | — | |
| Seed-1.5-VL (Thinking)Thinking Mode=true2026.01 | 68.7 | — | — | — | — | — | |
| GLM-4.6VParameters=106B-A12B2026.01 | 63.5 | — | — | — | — | — | |
| Claude-3.7-SonnetModel Category=Closed-Source2026.03 | 58.6 | — | — | — | — | — | |
| InternVL3.5Openness=semi-open, Model Size=8B2025.10 | 56.8 | — | — | — | — | — | |
| Bee-8BOpenness=fully open, Model Size=8B, Training Strategy=RL2025.10 | 50 | — | — | — | — | — | |
| Bee-8BOpenness=fully open, Model Size=8B, Training Strategy=SFT2025.10 | 46.8 | — | — | — | — | — | |
| Claude-3.5-Sonnet2026.02 | 46.48 | — | — | — | — | — | |
| Keye-VLOpenness=semi-open, Model Size=8B2025.10 | 46 | — | — | — | — | — | |
| InternVL3.5-2BThinking Mode=true2025.12 | 42.8 | — | — | — | — | — | |
| Claude 3.7Tool Use=false, Param Size=-2026.03 | 41.3 | — | — | — | — | — | |
| Qwen2.5-VL-32BParameters=32B2026.02 | 38.4 | — | — | — | — | — | |
| AVAR-ThinkerModel Category=Our model2026.03 | 37.4 | — | — | — | — | — | |
| V-STARData Size=40k2026.04 | 33.7 | — | — | — | — | — | |
| VL-Rethinker-7B + LEADBackbone=VL-Rethinker-7B, Decoding Strategy=LEAD2026.03 | 33.1 | — | — | — | — | — | |
| ThinkLite-VLModel Category=Multimodal Reasoning Models2026.03 | 32.9 | — | — | — | — | — | |
| ThinkLite-VLData Size=11k2026.04 | 32.9 | — | — | — | — | — | |
| AStarBackbone=Qwen2.5-7B, Training-free=true2025.02 | 32.7 | — | — | — | 39.4 | — | |
| R1-Onevision-7B + LEADBackbone=R1-Onevision-7B, Decoding Strategy=LEAD2026.03 | 32.4 | — | — | — | — | — | |
| VL-Cogito-7B + LEADBackbone=VL-Cogito-7B, Decoding Strategy=LEAD2026.03 | 32.4 | — | — | — | — | — | |
| VL-Rethinker-7BBackbone=VL-Rethinker-7B2026.03 | 32.3 | — | — | — | — | — | |
| VL-RethinkerData Size=39k2026.04 | 32.3 | — | — | — | — | — | |
| InternVL2.5-38BParameters=38B2026.02 | 32.2 | — | — | — | — | — | |
| VAPO-Thinker-7BModel Category=Our models, Parameter Scale=7B2025.09 | 31.9 | — | — | — | — | — | |
| SketchThinker-R1-7BMethod Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 31.7 | 65.5 | 0.484 | — | — | — | |
| Qwen3-VL-2B2025.12 | 31.6 | — | — | — | — | — | |
| GPT-4oModel Category=Closed-Source2026.03 | 31.2 | — | — | — | — | — | |
| Vanilla-R1Method Category=Direct Inference, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 31 | 221.1 | 0.14 | — | — | — | |
| VL-Cogito-7BBackbone=VL-Cogito-7B2026.03 | 30.7 | — | — | — | — | — | |
| VL-CogitoData Size=80k2026.04 | 30.7 | — | — | — | — | — | |
| GPT-4o2025.02 | 30.4 | — | — | — | 29.4 | — | |
| GPT-4oTool Use=false, Param Size=-2026.03 | 30.4 | — | — | — | — | — | |
| GPT-4oParadigm=Zero-Shot VLMs, Data Size=-2026.04 | 30.39 | — | — | — | — | — | |
| GPT-4o2026.02 | 30.3 | — | — | — | — | — | |
| MaLoRAModel=Qwen3-VL-8B, Training examples=2.8k2025.10 | 30.19 | — | — | — | — | — | |
| R1-OneVisionModel Category=Multimodal Reasoning Models2026.03 | 29.9 | — | — | — | — | — | |
| R1-Onevision-7BBackbone=R1-Onevision-7B2026.03 | 29.9 | — | — | — | — | — | |
| R1-OneVision-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 29.9 | — | — | — | — | — | |
| R1-OnevisionData Size=155k2026.04 | 29.9 | — | — | — | — | — | |
| Vision-R1-7B + LEADBackbone=Vision-R1-7B, Decoding Strategy=LEAD2026.03 | 29.7 | — | — | — | — | — | |
| ThinkPruneMethod Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 29.6 | 136.3 | 0.217 | — | — | — | |
| VL-Rethinker-7BParameters=7B2026.02 | 29.57 | — | — | — | — | — | |
| L1Method Category=Reinforcement-Learning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 29.5 | 146.7 | 0.201 | — | — | — | |
| InternVL3Tool Use=false, Param Size=8B2026.03 | 29.3 | — | — | — | — | — | |
| VL-RethinkerParadigm=Tool-use & RL Enhanced Reasoning, Data Size=39K2026.04 | 29.3 | — | — | — | — | — | |
| VeriThinkerMethod Category=Supervised-fine-tuning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 29.1 | 152.5 | 0.191 | — | — | — | |
| DeepEyesV2Tool Use=true, Param Size=7B2026.03 | 28.9 | — | — | — | — | — | |
| RuCL2026.02 | 28.88 | — | — | — | — | — | |
| C3oTMethod Category=Supervised-fine-tuning-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 28.8 | 125.5 | 0.229 | — | — | — | |
| InternVL3.5-8BParadigm=Zero-Shot VLMs, Data Size=70K2026.04 | 28.3 | — | — | — | — | — | |
| MIRROR(ours)Param Size=7B2026.02 | 28.29 | — | — | — | — | — | |
| Vision-R1-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 28.2 | — | — | — | — | — | |
| Auxiliary Reward2026.03 | 28.06 | — | — | — | — | — | |
| Perception-R1-7BParameters=7B2026.02 | 28.06 | — | — | — | — | — | |
| GeoFocusModel Scale=7B2026.02 | 28 | — | — | — | — | — | |
| AStarBackbone=Qwen2-VL-7B, Training-free=true2025.02 | 27.9 | — | — | — | 29.4 | — | |
| GRPOModel Scale=7B2026.02 | 27.7 | — | — | — | — | — | |
| MM-Eureka-7BParameters=7B2026.02 | 27.7 | — | — | — | — | — | |
| Semantic-backTool Use=false, Param Size=7B2026.03 | 27.7 | — | — | — | — | — | |
| ThymeTool Use=true, Param Size=7B2026.03 | 27.6 | — | — | — | — | — | |
| Trajectory-Hint2026.03 | 27.43 | — | — | — | — | — | |
| Chain-of-DraftMethod Category=Prompt-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 27.4 | 85.4 | 0.321 | — | — | — | |
| MaLoRAModel=Qwen2.5-VL-7B, Training examples=2.8k2025.10 | 27.35 | — | — | — | — | — | |
| LoRAModel=Qwen3-VL-8B, Training examples=2.8k2025.10 | 27.35 | — | — | — | — | — | |
| MIRROR(w/o tool)Param Size=7B2026.02 | 27.3 | — | — | — | — | — | |
| Vision-R1-7BParameters=7B2026.02 | 27.24 | — | — | — | — | — | |
| ThinkLite-VL-7BParameters=7B2026.02 | 27.24 | — | — | — | — | — | |
| Vision-R1-7BBackbone=Vision-R1-7B2026.03 | 27.2 | — | — | — | — | — | |
| Vision-R1Data Size=210k2026.04 | 27.2 | — | — | — | — | — | |
| BaseModel=Qwen3-VL-8B, Training examples=2.8k2025.10 | 27.2 | — | — | — | — | — | |
| Ctrl-RPower scaling factor (β)=0.22026.03 | 27.14 | — | — | — | — | — | |
| R1-VL-7B2025.02 | 27.1 | — | — | — | 23.6 | — | |
| MM-Eureka-7BModel Category=Multimodal Reasoning Models2026.03 | 26.9 | — | — | — | — | — | |
| Vision-SR1Model Category=Multimodal Reasoning Models2026.03 | 26.7 | — | — | — | — | — | |
| DeepEyesParam Size=7B2025.05 | 26.6 | — | — | — | — | — | |
| DeepEyesParadigm=Tool-use & RL Enhanced Reasoning, Data Size=47K2026.04 | 26.6 | — | — | — | — | — | |
| InternVL3.5-2BThinking Mode=false, Evaluation Framework=VLMEvalKit2025.12 | 26.5 | — | — | — | — | — | |
| VRETool Use=false, Param Size=7B2026.03 | 26.5 | — | — | — | — | — | |
| VLAA-Thinker-7BModel Category=Multimodal Reasoning Models2026.03 | 26.4 | — | — | — | — | — | |
| VLAA-Thinker-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 26.4 | — | — | — | — | — | |
| Constrained CoTMethod Category=Prompt-based, Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 26.2 | 79.2 | 0.331 | — | — | — | |
| URSA-8B2025.02 | 26.2 | — | — | — | 24.8 | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 25.9 | — | — | — | — | — | |
| OpenVLThinkerModel Category=Multimodal Reasoning Models2026.03 | 25.9 | — | — | — | — | — | |
| OpenVLThinkerData Size=59.2k2026.04 | 25.9 | — | — | — | — | — | |
| SCF-VRParadigm=Latent Reasoning, Data Size=30K, Backbone=Qwen2.5-VL-7B2026.04 | 25.89 | — | — | — | — | — | |
| OpenVLThinker2026.03 | 25.79 | — | — | — | — | — | |
| Qwen2.5-VL-3BParam Size=3B2026.02 | 25.66 | — | — | — | — | — | |
| Qwen2.5-VL*Param Size=7B2025.05 | 25.6 | — | — | — | — | — | |
| InternVL2.5-8B*Model Category=Open-source, Parameter Scale=8B, Reference Source=OpenCompass leaderboard2025.09 | 25.6 | — | — | — | — | — | |
| Qwen2.5-VL-7BParadigm=Zero-Shot VLMs, Data Size=-2026.04 | 25.6 | — | — | — | — | — | |
| Qwen2.5-VLOpenness=semi-open, Model Size=7B2025.10 | 25.4 | — | — | — | — | — | |
| GRPO2026.03 | 25.39 | — | — | — | — | — |