Multimodal Reasoning on MMstar (Pass@1 accuracy)
67.1Pass@1 AccuracyQwen2.5-VL-7B-Instruct + RFT
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-7B-Instruct + RFTModel Source Category=Ours, Training Strategy=RFT2026.04 | 67.1 | |
| InternVL2.5-26BModel Source Category=Open-Source Models2026.04 | 66.5 | |
| InternVL3-9BModel Source Category=Open-Source Models2026.04 | 66.3 | |
| Qwen2.5-VL-7B-Instruct + cold startModel Source Category=Ours, Training Strategy=cold start2026.04 | 66.2 | |
| LLaVA-OneVision-72BModel Source Category=Open-Source Models2026.04 | 65.8 | |
| Claude-3.5-SonnetModel Source Category=Closed-Source Models2026.04 | 65.1 | |
| GPT-4o-20240513Model Source Category=Closed-Source Models2026.04 | 64.7 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model2026.04 | 63.9 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Ours2026.04 | 63.9 | |
| InternVL2.5-8BModel Source Category=Open-Source Models2026.04 | 62.8 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model, Chain-of-Thought (CoT)=true2026.04 | 61.8 | |
| Qwen2-VL-7BModel Source Category=Open-Source Models2026.04 | 60.7 | |
| Gemini-1.5-ProModel Source Category=Closed-Source Models2026.04 | 59.1 | |
| MiniCPM-V2.6Model Source Category=Open-Source Models2026.04 | 57.5 | |
| GPT-4VModel Source Category=Closed-Source Models2026.04 | 56 | |
| Cambrian-34BModel Source Category=Open-Source Models2026.04 | 54.2 |