Multimodal Mathematical Reasoning on MathVerse (Pass@1 Accuracy)
47.2Pass@1 AccuracyQwen2.5-VL-7B-Instruct + RFT
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-7B-Instruct + RFTModel Source Category=Ours, Training Strategy=RFT2026.04 | 47.2 | |
| Claude-3.5-SonnetModel Source Category=Closed-Source Models2026.04 | 44.2 | |
| Qwen2.5-VL-7B-Instruct + cold startModel Source Category=Ours, Training Strategy=cold start2026.04 | 44.1 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model, Chain-of-Thought (CoT)=true2026.04 | 43.9 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model2026.04 | 43.3 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Ours2026.04 | 43.3 | |
| GPT-4o-20240513Model Source Category=Closed-Source Models2026.04 | 40.6 | |
| InternVL3-9BModel Source Category=Open-Source Models2026.04 | 35.3 | |
| Gemini-1.5-ProModel Source Category=Closed-Source Models2026.04 | 30.1 | |
| Qwen2-VL-7BModel Source Category=Open-Source Models2026.04 | 30.1 | |
| LLaVA-OneVision-72BModel Source Category=Open-Source Models2026.04 | 27.2 | |
| InternVL2.5-26BModel Source Category=Open-Source Models2026.04 | 24 | |
| InternVL2.5-8BModel Source Category=Open-Source Models2026.04 | 22.8 | |
| MiniCPM-V2.6Model Source Category=Open-Source Models2026.04 | 18.9 |