Medical Visual Question Answering on SLAKE (test)
91.8Closed AccuracyOurs
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OursModel Scale=13B2025.05 | 91.8 | 90.5 | 91 | — | — | — | — | — | — | — | — | — | — | |
| Dragonfly-Med2024.06 | 91.6 | — | — | 89.3 | — | — | — | — | — | — | — | — | — | |
| Yuan et al., 20232024.06 | 91.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| M2I22025.05 | 91.1 | 74.7 | 81.2 | — | — | — | — | — | — | — | — | — | — | |
| OursModel Scale=7B2025.05 | 90.4 | 90.1 | 90.2 | — | — | — | — | — | — | — | — | — | — | |
| BIOMEDCLIPModel Type=Representative non-VLM2026.04 | 89.7 | — | — | — | — | — | — | — | — | 82.05 | — | — | — | |
| parameter-efficient multi-level prompt frameworkParams=104M2026.05 | 89.7 | — | — | — | — | — | — | — | — | — | — | — | 87.3 | |
| PMC-CLIP2025.05 | 88 | 81.9 | 84.3 | — | — | — | — | — | — | — | — | — | — | |
| M3AEvision encoder=CLIP-ViT-B, language encoder=RoBERTa-base2022.09 | 87.82 | 80.31 | 83.25 | — | — | — | — | — | — | — | — | — | — | |
| M3AE2025.05 | 87.8 | 80.3 | 83.3 | — | — | — | — | — | — | — | — | — | — | |
| MedVInT-TE2025.05 | 87.7 | 88.2 | 88 | — | — | — | — | — | — | — | — | — | — | |
| DeRS-SMMoE Model=Med-MoE-Phi, Added Params=5.18M2025.03 | 87.2 | 84 | — | — | — | — | — | — | — | — | — | — | — | |
| DeRS-LMMoE Model=Med-MoE-Phi, Added Params=9.18M2025.03 | 86.5 | 84.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen 2.5 VL (Ours, SFT)Protocol=Supervised fine-tuned2026.04 | 86.4 | — | — | — | — | — | — | — | — | 85.5 | — | — | — | |
| DeRS-SMMoE Model=Med-MoE-Phi [20], Added Params.=0.84B+1.11M2025.03 | 86.3 | 84.3 | — | — | — | — | — | — | — | — | — | — | — | |
| MedVInT-TD2025.05 | 86.3 | 84.5 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| VanillaMoE Model=Med-MoE-Phi, Added Params=3.36B2025.03 | 85.8 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaMoE Model=Med-MoE-Phi [20], Added Params.=0.84B+2.52B2025.03 | 85.8 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | |
| DeRS-LMMoE Model=Med-MoE-Phi [20], Added Params.=0.84B+2.42M2025.03 | 85.6 | 84.3 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-MedProtocol=Supervised fine-tuned2026.04 | 85.34 | — | — | — | — | — | — | — | — | 83.08 | — | — | — | |
| VanillaMoE Model=Med-MoE-StableLM, Added Params=1.66B2025.03 | 85.3 | 82.4 | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaMoE Model=Med-MoE-StableLM [20], Added Params.=0.42B+1.24B2025.03 | 85.3 | 82.4 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-MedParams=7B2026.05 | 85.3 | — | — | — | — | — | — | — | — | — | — | — | 83 | |
| DeRS-SMMoE Model=Med-MoE-StableLM [20], Added Params.=0.42B+0.29M2025.03 | 84.9 | 83.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Med-Gemini2024.06 | 84.8 | — | — | 75.8 | — | — | — | — | — | — | — | — | — | |
| LLaVAParams=7B2026.05 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | 83.7 | |
| DeRS-SMMoE Model=Med-MoE-StableLM, Added Params=2.17M2025.03 | 84.4 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | |
| DeRS-LMMoE Model=Med-MoE-StableLM, Added Params=5.63M2025.03 | 84.4 | 83.6 | — | — | — | — | — | — | — | — | — | — | — | |
| DeRS-LMMoE Model=Med-MoE-StableLM [20], Added Params.=0.42B+1.20M2025.03 | 84.1 | 83.7 | — | — | — | — | — | — | — | — | — | — | — | |
| CPRD-BAN2022.09 | 83.4 | 79.5 | 81.1 | — | — | — | — | — | — | — | — | — | — | |
| CPRD-BAN2025.05 | 83.4 | 79.5 | 81.1 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Med2024.06 | 83.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Med2025.05 | 83.2 | 84.7 | 84.1 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-TriParams=8B2026.05 | 83.2 | — | — | — | — | — | — | — | — | — | — | — | 82.3 | |
| BiMediX2-8B2025.03 | 83.1 | 72.9 | — | — | 77.8 | 78.6 | — | — | — | — | — | — | — | |
| PeFoMedParams=33M, Configuration=LoRA2026.05 | 82.7 | — | — | — | — | — | — | — | — | — | — | — | 81.3 | |
| PUBMEDCLIPModel Type=Representative non-VLM2026.04 | 82.5 | — | — | — | — | — | — | — | — | 78.4 | — | — | — | |
| DPOBackbone=Qwen3-VL-4B-Instruct, Text* Pref.=×, Img. Pref.=×2026.06 | 82.41 | 50.78 | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP-ViT2025.05 | 82.1 | 84.3 | 83.3 | — | — | — | — | — | — | — | — | — | — | |
| MASK-DPOBackbone=HuatuoGPT-Vision-7B, Text* Pref.=✓, Img. Pref.=×2026.06 | 81.32 | 57.1 | — | — | — | — | — | — | — | — | — | — | — | |
| FIRE-MPOBackbone=HuatuoGPT-Vision-7B, Text* Pref.=✓, Img. Pref.=✓2026.06 | 81.32 | 61.26 | — | — | — | — | — | — | — | — | — | — | — | |
| FIRE-MPOBackbone=Qwen3-VL-4B-Instruct, Text* Pref.=✓, Img. Pref.=✓2026.06 | 81.32 | 60.25 | — | — | — | — | — | — | — | — | — | — | — | |
| MASK-DPOBackbone=Qwen3-VL-4B-Instruct, Text* Pref.=✓, Img. Pref.=×2026.06 | 80.49 | 59.54 | — | — | — | — | — | — | — | — | — | — | — | |
| MEVF-BAN2022.09 | 79.8 | 77.8 | 78.6 | — | — | — | — | — | — | — | — | — | — | |
| MEVF-BAN2025.05 | 79.8 | 77.8 | 78.6 | — | — | — | — | — | — | — | — | — | — | |
| DPOBackbone=HuatuoGPT-Vision-7B, Text* Pref.=×, Img. Pref.=×2026.06 | 79.4 | 55.38 | — | — | — | — | — | — | — | — | — | — | — | |
| RRPOBackbone=HuatuoGPT-Vision-7B, Text* Pref.=✓, Img. Pref.=×2026.06 | 79.12 | 59.11 | — | — | — | — | — | — | — | — | — | — | — | |
| SAN2022.09 | 79.1 | 74 | 76 | — | — | — | — | — | — | — | — | — | — | |
| BAN2022.09 | 79.1 | 74.6 | 76.3 | — | — | — | — | — | — | — | — | — | — | |
| MEVF-SAN2022.09 | 78.4 | 75.3 | 76.5 | — | — | — | — | — | — | — | — | — | — | |
| mDPOBackbone=Qwen3-VL-4B-Instruct, Text* Pref.=×, Img. Pref.=✓2026.06 | 77.19 | 56.24 | — | — | — | — | — | — | — | — | — | — | — | |
| RRPOBackbone=Qwen3-VL-4B-Instruct, Text* Pref.=✓, Img. Pref.=×2026.06 | 77.19 | 60.11 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4B-InstructText* Pref.=-, Img. Pref.=-2026.06 | 76.64 | 56.24 | — | — | — | — | — | — | — | — | — | — | — | |
| BioD2C2025.03 | 76.3 | 74.2 | — | — | 76.6 | 81 | — | — | — | — | — | — | — | |
| RadFM2025.03 | 75.2 | 72.5 | — | — | 74.6 | 69.5 | — | — | — | — | — | — | — | |
| MFB2022.09 | 75 | 72.2 | 73.3 | — | — | — | — | — | — | — | — | — | — | |
| mDPOBackbone=HuatuoGPT-Vision-7B, Text* Pref.=×, Img. Pref.=✓2026.06 | 73.59 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | |
| MMedPOText* Pref.=×, Img. Pref.=×2026.06 | 73.08 | 53.99 | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-38BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 71.5 | — | — | — | — | — | — | — | — | 64.6 | — | — | — | |
| FiSAOText* Pref.=×, Img. Pref.=×2026.06 | 70.46 | 52.69 | — | — | — | — | — | — | — | — | — | — | — | |
| HuatuoGPT-Vision-7BText* Pref.=-, Img. Pref.=-2026.06 | 70.33 | 47.49 | — | — | — | — | — | — | — | — | — | — | — | |
| o3Protocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | 67.2 | — | — | — | — | — | — | — | — | 44.6 | — | — | — | |
| Qwen 2.5 VLProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | 66.9 | — | — | — | — | — | — | — | — | 51.2 | — | — | — | |
| GLM-4VProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 66.7 | — | — | — | — | — | — | — | — | 66.9 | — | — | — | |
| InternVL-6BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 66.3 | — | — | — | — | — | — | — | — | 62.9 | — | — | — | |
| InternVL-30BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 65.8 | — | — | — | — | — | — | — | — | 64.3 | — | — | — | |
| GPT-5Protocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | 64.1 | — | — | — | — | — | — | — | — | 43.2 | — | — | — | |
| LLaVAProtocol=Supervised fine-tuned2026.04 | 63.22 | — | — | — | — | — | — | — | — | 78.18 | — | — | — | |
| STLLaVA-MedText* Pref.=×, Img. Pref.=×2026.06 | 61.75 | 48.65 | — | — | — | — | — | — | — | — | — | — | — | |
| THYMEProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 59.6 | — | — | — | — | — | — | — | — | 46.1 | — | — | — | |
| VL ENCODER–DECODERModel Type=Representative non-VLM2026.04 | 58.29 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Q2ATRANSFORMERModel Type=Representative non-VLM2026.04 | 54.85 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Med-1.52025.03 | 53.6 | 33.4 | — | — | 0.2 | 41.3 | — | — | — | — | — | — | — | |
| MedVInT-TD2025.03 | 49.8 | 33.8 | — | — | 21.3 | 35.1 | — | — | — | — | — | — | — | |
| GLM-4.5VProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | 49.7 | — | — | — | — | — | — | — | — | 47.6 | — | — | — | |
| GPT-4VMode=zero-shot2026.05 | 43.6 | — | — | — | — | — | — | — | — | — | — | — | 33.6 | |
| PREFIX T. MEDICAL LMModel Type=Representative non-VLM2026.04 | 40 | — | — | — | — | — | — | — | — | 82.01 | — | — | — | |
| Gemini 2.5 ProProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | 39 | — | — | — | — | — | — | — | — | 23.5 | — | — | — | |
| MMADAProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 37.6 | — | — | — | — | — | — | — | — | 7 | — | — | — | |
| Qwen2.5-VL-7BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 37.1 | — | — | — | — | — | — | — | — | 53.3 | — | — | — | |
| BioMedGPT2025.03 | 24.8 | 25.9 | — | — | 17.5 | 26 | — | — | — | — | — | — | — | |
| LLaVAProtocol=Zero-shot, Model Type=open-source VLM2026.04 | 13.46 | — | — | — | — | — | — | — | — | 18.55 | — | — | — | |
| ARMed-RModel Type=Fine-tuned VLM, Protocol=Reasoner2025.08 | — | — | — | — | 76.14 | 76.65 | 99.46 | 98.13 | 87.6 | — | — | — | — | |
| Base Modelbase_model=Qwen2.5-VL-3B-Instruct2026.03 | — | — | 68.73 | — | — | 43.97 | — | — | — | 34.7 | — | — | — | |
| CheXagent2024.08 | — | — | 71.1 | — | — | — | — | — | — | 73.2 | — | — | — | |
| DeepEyes†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | — | — | 59.7 | — | — | — | — | — | — | — | — | — | — | |
| EN-INFbase_model=Qwen2.5-VL-3B-Instruct2026.03 | — | — | 65.92 | — | — | 45.23 | — | — | — | 35.55 | — | — | — | |
| Gemini-2.5-proFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | — | — | 72.7 | — | — | — | — | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework II, Inference Mode=Direct Inference2026.03 | — | — | 77.7 | — | — | — | — | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | — | — | 73.9 | — | — | — | — | — | — | — | — | — | — | |
| GMAI-VLCategory=Medical-Specific Models2025.11 | — | — | 71.9 | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oCategory=General Vision-Language Models2025.11 | — | — | 50.1 | — | — | — | — | — | — | — | — | — | — | |
| GPT-5Framework=Framework II, Inference Mode=Direct Inference2026.03 | — | — | 73.2 | — | — | — | — | — | — | — | — | — | — | |
| GRIT†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | — | — | 57.1 | — | — | — | — | — | — | — | — | — | — | |
| HuatuoGPT-V-7BModel Type=Medical VLM2025.08 | — | — | — | — | 13.11 | 20.25 | 88.23 | 86.26 | 51.96 | — | — | — | — | |
| InternVLfine-tuned=false, category=General Vision-Language Models, evaluation=5-fold cross validation2026.03 | — | — | 70.8 | — | — | — | — | — | — | — | 61.2 | 59.8 | — | |
| InternVL-2Category=General Vision-Language Models2025.11 | — | — | 62.7 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-14BModel Type=General VLM2025.08 | — | — | — | — | 16.87 | 22.32 | 88.62 | 87.02 | 53.71 | — | — | — | — | |
| InternVL3-2BModel Type=General VLM2025.08 | — | — | — | — | 39.18 | 43.49 | 93.59 | 91.67 | 66.98 | — | — | — | — | |
| InternVL3-8BModel Type=General VLM2025.08 | — | — | — | — | 17.36 | 23.62 | 87.29 | 85.38 | 53.41 | — | — | — | — |