Mathematical Reasoning on MathVision (test)
71.9AccuracyGPT5 mini
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT5 miniModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 71.9 | — | — | |
| Qwen3-VL 32BModel Source Category=Open-weight VLMs, Model Parameters=32B, Evaluation Mode=Thinking Mode2026.01 | 70.2 | — | — | |
| MMFineReason-8BModel Source Category=Ours, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 67.1 | — | — | |
| Qwen3-VL 30B-A3BModel Source Category=Open-weight VLMs, Model Parameters=30B-A3B, Evaluation Mode=Thinking Mode2026.01 | 65.7 | — | — | |
| Gemini-2.5 FlashModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 64.3 | — | — | |
| Qwen3-VL 8BModel Source Category=Open-weight VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 62.7 | — | — | |
| MMFineReason-4BModel Source Category=Ours, Model Parameters=4B, Evaluation Mode=Thinking Mode2026.01 | 61.3 | — | — | |
| MiMo-VL-7B-SFT-25082025.12 | 57.2 | — | — | |
| MiMo-VL-Miloco-7B2025.12 | 54 | — | — | |
| MMR1 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 48.4 | — | — | |
| MMFineReason-2BModel Source Category=Ours, Model Parameters=2B, Evaluation Mode=Thinking Mode2026.01 | 45.3 | — | — | |
| Qwen2.5-VL-32B + AT-RL (Ours)Zero-shot=true2026.02 | 44.2 | — | — | |
| OMR 7BModel Source Category=Open-source VLMs, Model Parameters=7B, Evaluation Mode=Thinking Mode2026.01 | 43.6 | — | — | |
| Qwen2.5-VL-72B InstructZero-shot=true2026.02 | 43.1 | — | — | |
| Ours (CDRL + CA-TTS)Framework Type=Training-Based, Base Model=Qwen2.5-VL-7B2026.03 | 42.4 | 38.4 | 45.6 | |
| Qwen2.5-VL-32B + VPPOZero-shot=true2026.02 | 42.3 | — | — | |
| Gemini 2.0 FlashZero-shot=true2026.02 | 41.3 | — | — | |
| Qwen2.5-VL-32B InstructZero-shot=true2026.02 | 38.2 | — | — | |
| HoneyBee 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 37.4 | — | — | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 34.64 | — | — | |
| Claude 3.5 SonnetZero-shot=true2026.02 | 33.5 | — | — | |
| MiMo-VLtraining=7B-SFT, parameters=7B2026.01 | 32.99 | — | — | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 31.32 | — | — | |
| VL-RethinkerFramework Type=Training-Based, Base Model=Qwen2.5-VL-7B2026.03 | 30.7 | 22.1 | 39 | |
| OpenAI GPT-4oZero-shot=true2026.02 | 30.6 | — | — | |
| Qwen3-VLparameters=8B2026.01 | 30.56 | — | — | |
| Majority VotingFramework Type=Training-Free, Base Model=Qwen2.5-VL-7B2026.03 | 30.1 | 26.2 | 33.2 | |
| R1-OnevisionFramework Type=Training-Based, Base Model=Qwen2.5-VL-7B2026.03 | 29.9 | 20.8 | 37.8 | |
| MiMo-VLtraining=7B-RL, parameters=7B2026.01 | 29.77 | — | — | |
| We-ThinkFramework Type=Training-Based, Base Model=Qwen2.5-VL-7B2026.03 | 29.7 | 20.9 | 37.4 | |
| DeepconfFramework Type=Training-Free, Base Model=Qwen2.5-VL-7B2026.03 | 29.6 | 26.4 | 32.3 | |
| InternVL3.5parameters=8B2026.01 | 27.11 | — | — | |
| RTWIModel Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 25.5 | — | — | |
| Qwen2.5-VL-7B + Cont. Rewardreward_type=Continuous2025.11 | 24.81 | — | — | |
| SCModel Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 24.4 | — | — | |
| LLaVA-OVversion=1.5, parameters=8B2026.01 | 24.11 | — | — | |
| Self-Cer.Model Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 24 | — | — | |
| Vision-Zeroexternal supervision=true2025.11 | 23.96 | — | — | |
| Qwen2.5-VL-7B (Baseline)2025.11 | 23.91 | — | — | |
| MiniCPM-Vversion=4.5, parameters=8B2026.01 | 23.72 | — | — | |
| RTWIModel Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 23.4 | — | — | |
| Pass@1Framework Type=Training-Free, Base Model=Qwen2.5-VL-7B2026.03 | 23 | 24.2 | 21.7 | |
| Intern-S1variant=mini, parameters=9B2026.01 | 22.76 | — | — | |
| CISCModel Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 22.7 | — | — | |
| DeepconfModel Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 22.7 | — | — | |
| Qwen2.5-VL-7B + Discrete Rewardreward_type=Discrete2025.11 | 22.52 | — | — | |
| DreamPRMFramework Type=Training-Based, Base Model=InternVL-2.5-8B2026.03 | 22.1 | — | — | |
| BaseModel Variant=Qwen3-VL Instruct, Setting=Offline2026.02 | 21.4 | — | — | |
| CISCModel Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 20.1 | — | — | |
| SCModel Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 19.8 | — | — | |
| DeepconfModel Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 19.7 | — | — | |
| BaseModel Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 18.7 | — | — | |
| Self-Cer.Model Variant=Qwen3-VL Thinking, Setting=Offline2026.02 | 18 | — | — |