Multimodal Mathematical Reasoning on MathVista (Pass@1 accuracy)
73.8Pass@1 AccuracyQwen2.5-VL-7B-Instruct + RFT
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-7B-Instruct + RFTModel Source Category=Ours, Training Strategy=RFT2026.04 | 73.8 | |
| InternVL3-9BModel Source Category=Open-Source Models2026.04 | 71.5 | |
| Qwen2.5-VL-7B-Instruct + cold startModel Source Category=Ours, Training Strategy=cold start2026.04 | 71.1 | |
| Gemini-1.5-ProModel Source Category=Closed-Source Models2026.04 | 68.7 | |
| InternVL2.5-26BModel Source Category=Open-Source Models2026.04 | 68.2 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model, Chain-of-Thought (CoT)=true2026.04 | 68.2 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model2026.04 | 68.1 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Ours2026.04 | 68.1 | |
| LLaVA-OneVision-72BModel Source Category=Open-Source Models2026.04 | 67.1 | |
| Claude-3.5-SonnetModel Source Category=Closed-Source Models2026.04 | 64.7 | |
| InternVL2.5-8BModel Source Category=Open-Source Models2026.04 | 64.5 | |
| Qwen2-VL-7BModel Source Category=Open-Source Models2026.04 | 62.3 | |
| MiniCPM-V2.6Model Source Category=Open-Source Models2026.04 | 60.8 | |
| GPT-4o-20240513Model Source Category=Closed-Source Models2026.04 | 60 | |
| Cambrian-34BModel Source Category=Open-Source Models2026.04 | 53.2 |