Mathematical Reasoning on MathVision (Acc@1, Acc@4)
72Top-1 AccuracyGPT-5-Thinking*
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5-Thinking*Model Category=Closed-Source Models2026.04 | 72 | 79.2 | |
| Gemini-2.5-Pro*Model Category=Closed-Source Models2026.04 | 69.1 | 72.5 | |
| MUPO-Thinker-7BModel Category=Our Models2026.04 | 31.3 | 39.7 | |
| R1-OneVision-7BModel Category=Open-Source Models2026.04 | 29.9 | 34.1 | |
| GRPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | 29.3 | — | |
| Vision-R1-7BModel Category=Open-Source Models2026.04 | 28.2 | 32.9 | |
| MUPO-Thinker-3BModel Category=Our Models2026.04 | 27.8 | 35.4 | |
| GSPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | 27.6 | — | |
| GRPOModel Scale=Qwen2.5-VL-7B2026.06 | 26.6 | — | |
| SAPOModel Scale=Qwen2.5-VL-7B2026.06 | 26.6 | — | |
| GSPOModel Scale=Qwen2.5-VL-7B2026.06 | 26.6 | — | |
| VLAA-Thinker-7BModel Category=Open-Source Models2026.04 | 26.4 | 30.3 | |
| InternVL2.5-8B*Model Category=Open-Source Models2026.04 | 25.6 | 29.4 | |
| SAPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | 24.6 | — | |
| DAPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | 24.4 | — | |
| GSPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | 24 | — | |
| Zero-ShotModel Scale=Qwen2.5-VL-7B2026.06 | 23.7 | — | |
| Qwen2.5-VL-7BModel Category=Open-Source Models2026.04 | 23.2 | 41.6 | |
| Zero-ShotModel Scale=Qwen2.5-VL-3B2026.06 | 23 | — | |
| DAPOModel Scale=Qwen2.5-VL-7B2026.06 | 23 | — | |
| DAPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | 22.8 | — | |
| GRPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | 22.4 | — | |
| GRPOModel Scale=Qwen2.5-VL-3B2026.06 | 21.4 | — | |
| SAPOModel Scale=Qwen2.5-VL-3B2026.06 | 20.1 | — | |
| SAPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | 19.7 | — | |
| GSPOModel Scale=Qwen2.5-VL-3B2026.06 | 19.7 | — | |
| DAPOModel Scale=Qwen2.5-VL-3B2026.06 | 17.1 | — |