Mathematical Multimodal Reasoning on MM-Math
59.3AccuracyVision-R1-72B*
Evaluation Results
| Method | Links | |
|---|---|---|
| Vision-R1-72B*Params.=72B, additional data during RL training=true2025.03 | 59.3 | |
| Vision-R1-32B*Params.=32B, additional data during RL training=true2025.03 | 55.3 | |
| Qwen2.5-VL-72BParams.=72B2025.03 | 45.6 | |
| OursModel Category=Reasoning MLLMs, Training Strategy=DAPO2026.01 | 43.8 | |
| Ours [with DPS]Model Category=Reasoning MLLMs, Training Strategy=DPS2026.01 | 43.4 | |
| Ours [with DPS and annealing]Model Category=Reasoning MLLMs, Training Strategy=Two-stage RL2026.01 | 43.4 | |
| Vision-R1-7BParams.=7B2025.03 | 40.2 | |
| VisonR1 7BModel Category=Reasoning MLLMs2026.01 | 40 | |
| VLAA-Thinker 7BModel Category=Reasoning MLLMs2026.01 | 39 | |
| Qwen2.5VL 7BModel Category=Open-Source General MLLMs2026.01 | 36.4 | |
| MixedR1 7BModel Category=Reasoning MLLMs2026.01 | 35.8 | |
| Qwen2.5-VL-32BParams.=32B2025.03 | 34.9 | |
| Qwen2.5-VL-7BParams.=7B2025.03 | 34.1 | |
| R1-Onevision 7BModel Category=Reasoning MLLMs2026.01 | 32.9 | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 31.8 | |
| GPT-4oParams.=-2025.03 | 31.8 | |
| Vision-R1-LlamaV-CI-11BParameters=11B2025.03 | 26.1 | |
| MulberryModel Category=Reasoning MLLMs2026.01 | 23.7 | |
| Mulberry-7BParams.=7B2025.03 | 23.7 | |
| GPT-4VModel Category=Closed-Source MLLMs2026.01 | 23.1 | |
| GPT-4VParams.=-2025.03 | 23.1 | |
| Mulberry-Llama-11BParameters=11B2025.03 | 18.7 | |
| InternVL2.5-78BParams.=78B2025.03 | 17.8 | |
| LLaVA-Cot-11BParameters=11B2025.03 | 16.5 | |
| LLaVA-CoT-11BParams.=11B2025.03 | 16.5 | |
| Llama-3.2-11B-VParameters=11B2025.03 | 4.1 |