Multimodal Mathematical Reasoning on FlowVerse Text Limited
40.5CoT ErrorInternVL2.5-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL2.5-8BModel Type=Open-source MLLMs2025.03 | 40.5 | 38.4 | |
| InfiMM-Math-7BModel Type=Open-source MLLMs2025.03 | 40.6 | 36.7 | |
| Math-LLaVA-13BModel Type=Math-specialized MLLMs2025.03 | 44.4 | 37.4 | |
| Qwen-VL-MaxModel Type=Closed-source MLLMs2025.03 | 46.7 | 38.3 | |
| MultiMath-7BModel Type=Math-specialized MLLMs2025.03 | 49.9 | 42.9 | |
| SVE-Math-Qwen2.5-7BModel Type=Math-specialized MLLMs2025.03 | 53.4 | 45.8 | |
| Qwen2-VL-72BModel Type=Open-source MLLMs2025.03 | 54.3 | 45.7 | |
| VLM-R1-7B†Model Type=Open-source MLLMs2025.03 | 57.9 | 49.8 | |
| GPT-4o-miniModel Type=Closed-source MLLMs2025.03 | 58.2 | 53.2 | |
| Claude-3.5-SonnetModel Type=Closed-source MLLMs2025.03 | 58.7 | 50.3 | |
| GPT-4oModel Type=Closed-source MLLMs2025.03 | 58.7 | 54.4 | |
| Qwen2.5-VL-7BModel Type=Open-source MLLMs2025.03 | 58.9 | 51.3 | |
| MathFlowModel Type=Open-source MLLMs, Backbone=Qwen2.5-VL-7B2025.03 | 60.8 | 52.2 | |
| InternVL2.5-78BModel Type=Open-source MLLMs2025.03 | 64.1 | 60.3 | |
| GPT-4VModel Type=Closed-source MLLMs2025.03 | 65 | 55 | |
| Gemini-2.5-proModel Type=Closed-source MLLMs2025.03 | 66.1 | 60.8 | |
| GPT-5Model Type=Closed-source MLLMs2025.03 | 73.5 | 66.7 | |
| MathFlowModel Type=Closed-source MLLMs, Backbone=GPT-52025.03 | 73.8 | 67.2 |