Multimodal Reasoning Accuracy on MM-Vet
76.2Pass@1 AccuracyInternVL3-9B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL3-9BModel Source Category=Open-Source Models2026.04 | 76.2 | — | |
| Qwen2.5-VL-7B-Instruct + RFTModel Source Category=Ours, Training Strategy=RFT2026.04 | 73.8 | — | |
| Claude-3.5-SonnetModel Source Category=Closed-Source Models2026.04 | 70.1 | — | |
| GPT-4o-20240513Model Source Category=Closed-Source Models2026.04 | 69.1 | — | |
| GPT-4VModel Source Category=Closed-Source Models2026.04 | 67.5 | — | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model, Chain-of-Thought (CoT)=true2026.04 | 67.2 | — | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model2026.04 | 67.1 | — | |
| Qwen2.5-VL-7B-InstructModel Source Category=Ours2026.04 | 67.1 | — | |
| InternVL2.5-26BModel Source Category=Open-Source Models2026.04 | 65 | — | |
| Qwen2.5-VL-7B-Instruct + cold startModel Source Category=Ours, Training Strategy=cold start2026.04 | 64.2 | — | |
| Gemini-1.5-ProModel Source Category=Closed-Source Models2026.04 | 64 | — | |
| InternVL2.5-8BModel Source Category=Open-Source Models2026.04 | 62.8 | — | |
| Qwen2-VL-7BModel Source Category=Open-Source Models2026.04 | 62 | — | |
| LLaVA-OneVision-72BModel Source Category=Open-Source Models2026.04 | 60.6 | — | |
| MiniCPM-V2.6Model Source Category=Open-Source Models2026.04 | 60 | — | |
| Cambrian-34BModel Source Category=Open-Source Models2026.04 | 53.2 | — | |
| Imp-v1LM=Phi-2-2.7B, Res.=3842024.12 | — | 33.1 | |
| LLaVA-1.5LM=Vicuna-7B, Res.=3362024.12 | — | 30.5 | |
| LLaVA-PhiLM=Phi-2-2.7B, Res.=3362024.12 | — | 28.9 | |
| Mipha-3BLM=Phi-2-2.7B, Res.=3842024.12 | — | 32.1 | |
| MoE-LLaVA-3.6BLM=Phi-2-2.7B, Res.=3842024.12 | — | 35.9 | |
| mPLUG-Owl2LM=LLaMA-7B, Res.=4482024.12 | — | 36.2 | |
| OlympusLM=Phi-2-2.7B, Res.=3842024.12 | — | 33.8 | |
| TinyLLaVALM=Phi-2-2.7B, Res.=3842024.12 | — | 32 |