Medical Visual Question Answering on MMMU Health & Medicine (test)
74.5AccuracyOpenAI-o3
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenAI-o3Category=Close-Source SOTA2026.04 | 74.5 | |
| GPT-4.1Category=Close-Source SOTA2026.04 | 73.6 | |
| Gemini 2.5 ProCategory=Close-Source SOTA2026.04 | 72.8 | |
| GPT-5Category=Close-Source SOTA2026.04 | 70.7 | |
| InternVL3-8BCategory=Open-Source SOTA2026.04 | 62.3 | |
| Qwen2.5-VL-32BCategory=Open-Source SOTA2026.04 | 60.1 | |
| HuatuoGPT-Vision-34BCategory=Medical MLLMs2026.04 | 60.1 | |
| MedEyesBackbone=Qwen2.5-VL-3B, Visual Expert=MedPLIB, Resolution=336x336, Patch size=142025.11 | 59.7 | |
| MMedAgent-RL-7BCategory=Multimodal Medical Agents2026.04 | 58.9 | |
| PixelReasoner-RL-v1-7BCategory=MLLMs can Think with Images2026.04 | 58 | |
| DeepEyes-7BCategory=MLLMs can Think with Images2026.04 | 57.8 | |
| Mini-o3-7B-v1Category=MLLMs can Think with Images2026.04 | 57.4 | |
| VILA-M3-40BCategory=Multimodal Medical Agents2026.04 | 56.6 | |
| MedLVRCategory=MLLMs can Think with Images2026.04 | 56.6 | |
| MedAgent-ProCategory=Multimodal Medical Agents2026.04 | 52.9 | |
| InternVL-2Category=General Vision-Language Models2025.11 | 52.7 | |
| GMAI-VLCategory=Medical-Specific Models2025.11 | 51.2 | |
| GRIT†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 49.5 | |
| AURACategory=Multimodal Medical Agents2026.04 | 49.3 | |
| DeepEyes†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 49.1 | |
| Med-FlamingoCategory=Medical-Specific Models2025.11 | 47.5 | |
| Qwen2.5-VL-7BCategory=MLLMs can Think with Images2026.04 | 46.4 | |
| MedVLM-R1Category=MLLMs can Think about Images2026.04 | 45.9 | |
| Med-R1Category=MLLMs can Think about Images2026.04 | 44.7 | |
| MMedAgent-7BCategory=Multimodal Medical Agents2026.04 | 44.1 | |
| MEDVISTAGYMBackbone=Qwen3vl-8B, Category=MLLMs can Think with Images2026.04 | 42.9 | |
| LLaVA-Next-13BCategory=Open-Source SOTA2026.04 | 40.1 | |
| SMR-AgentsCategory=Multimodal Medical Agents2026.04 | 40.1 | |
| LLaVA-Med-7BCategory=Medical MLLMs2026.04 | 38.8 | |
| LLaVA-v1.5-8BCategory=Open-Source SOTA2026.04 | 38.2 | |
| LLaVA-MedCategory=Medical-Specific Models2025.11 | 36.9 | |
| MedVLM-R1Category=Reinforcement Learning Methods2025.11 | 35.5 | |
| Qwen2.5-VL-3BCategory=General Vision-Language Models2025.11 | 34.3 | |
| LLaVA-Next-7BCategory=Open-Source SOTA2026.04 | 33.1 | |
| Med-R1Category=Reinforcement Learning Methods2025.11 | 32.7 | |
| MedVInTCategory=Medical-Specific Models2025.11 | 28.3 | |
| Med-FlamingoCategory=Medical MLLMs2026.04 | 28.3 | |
| RadFMCategory=Medical-Specific Models2025.11 | 27 | |
| RadFMCategory=Medical MLLMs2026.04 | 27 |