Medical Visual Question Answering on PMC-VQA
65.8AccuracyGPT-5.4
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5.42026.05 | 65.8 | |
| Claude-Opus-4.62026.05 | 64.8 | |
| Rform + Rexact + RDTWReward functions=Rform + Rexact + RDTW2026.04 | 64.1 | |
| Ours-7BModel Category=Medical Models2026.04 | 63.75 | |
| GPT-5Model size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 62 | |
| Gemini-2.5-proModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 62 | |
| VITAL2026.05 | 61.7 | |
| OpenAI-o3Category=Close-Source SOTA2026.04 | 61.5 | |
| Rform + RexactReward functions=Rform + Rexact2026.04 | 61.3 | |
| Fleming-VL2026.05 | 61.3 | |
| GPT-5Category=Close-Source SOTA2026.04 | 60.2 | |
| GPT-5-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 60 | |
| Claude-4.5-haikuModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 59 | |
| Claude-4.5-sonnetModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 59 | |
| GEMINI-3-FLASHModel Type=Proprietary2026.01 | 58.1 | |
| MEDVISTA-R1Model size=7-13B parameters2026.01 | 58 | |
| GPT-5Model Type=Proprietary2026.01 | 57.7 | |
| LINGSHU-7B + MED-SCOUTParameters=7B, Enhancement=Med-Scout2026.01 | 57.4 | |
| InternVL3-8BModel size=7-13B parameters2026.01 | 56.5 | |
| MedAgentsBase Model=GPT-4o2025.08 | 56.5 | |
| TMA-AllComponBase Model=GPT-4o2025.08 | 56.4 | |
| MDAgentsBase Model=GPT-4o2025.08 | 56.4 | |
| MedTutor-R1Backbone=Qwen2.5VL-7B-Instruct2025.12 | 56.3 | |
| LINGSHU-7BParameters=7B2026.01 | 56.3 | |
| Lingshu-7BModel Category=Medical Models2026.04 | 56.3 | |
| MEDIC-ADSize=7B2026.03 | 56.1 | |
| GPT-o4-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 56 | |
| Cold-start w/o ReasoningModel size=7-13B parameters, Reasoning-stage setting=removed, SFT-stage setting=Cold-start2026.01 | 56 | |
| LingshuSize=7B2026.03 | 56 | |
| DyLANBase Model=GPT-4o2025.08 | 56 | |
| ReConcileBase Model=GPT-4o2025.08 | 56 | |
| HUATUOGPT-VISION-7B + MED-SCOUTParameters=7B, Enhancement=Med-Scout2026.01 | 55.9 | |
| Gemini 2.5 ProCategory=Close-Source SOTA2026.04 | 55.9 | |
| RformReward functions=Rform2026.04 | 55.8 | |
| Qwen3-VL-8BThinking=true2026.05 | 55.75 | |
| Citrus-VSize=8B2026.03 | 55.6 | |
| GPT-4.1Category=Close-Source SOTA2026.04 | 55.2 | |
| Qwen3-VL-8BZero-shot=true2026.05 | 54.65 | |
| SFTTraining configuration=SFT2026.04 | 54.5 | |
| InternVL3-8BModel Category=General Models2026.04 | 53.95 | |
| Qwen2.5-VL-7BModel Category=General Models2026.04 | 53.65 | |
| MedLVRCategory=MLLMs can Think with Images2026.04 | 53.6 | |
| MedLVR2026.05 | 53.6 | |
| HuatuoGPT-V-7BModel Category=Medical Models2026.04 | 53.3 | |
| HuatuoGPT-Vision2026.05 | 53.2 | |
| MedTutor-R1 w/ LLaVA-basedBackbone=LLaVA2025.12 | 53.09 | |
| HUATUOGPT-VISION-7BParameters=7B2026.01 | 53 | |
| InternVL3-8BCategory=Open-Source SOTA2026.04 | 52.7 | |
| Qwen3vl-8BModel size=7-13B parameters2026.01 | 52.5 | |
| MedTutor-R1 w/o RLReinforcement Learning=disabled2025.12 | 52.28 | |
| INTERNVL3-8BParameters=8B2026.01 | 52 | |
| QWEN2.5-VL-7B-INSTRUCTParameters=7B2026.01 | 51.8 | |
| LVRWeights=Fine-tuned2026.05 | 51.56 | |
| GRPO w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, Reasoning-stage setting=GRPO2026.01 | 51.5 | |
| HuatuoGPT-Vision-34BCategory=Medical MLLMs2026.04 | 51.4 | |
| SIM-CoTBackbone=Qwen3-VL-8B + LoRA2026.05 | 51.3 | |
| Cold-start w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, SFT-stage setting=Cold-start2026.01 | 51 | |
| Qwen2.5-VL-32BCategory=Open-Source SOTA2026.04 | 50.4 | |
| QWEN2.5-VL-3B-INSTRUCTParameters=3B2026.01 | 50.2 | |
| CoconutBackbone=Qwen3-VL-8B + LoRA2026.05 | 50.2 | |
| MEDVISTAGYM (Qwen3vl-8B)Model size=7-13B parameters, Tool-access setting=enabled2026.01 | 49.5 | |
| MEDVISTA-R1Model size=< 7B parameters2026.01 | 49.5 | |
| MEDVISTAGYMBackbone=Qwen3vl-8B, Category=MLLMs can Think with Images2026.04 | 49.5 | |
| Qwen2.5vl-7BModel size=7-13B parameters2026.01 | 49 | |
| Direct GRPO w/o cold-startModel size=7-13B parameters, SFT-stage setting=removed, Reasoning-stage setting=Direct GRPO2026.01 | 49 | |
| Qwen2.5-VL-7BCategory=MLLMs can Think with Images2026.04 | 49 | |
| MCOUTWeights=Fine-tuned2026.05 | 48.89 | |
| MEDGEMMA-4B-ITParameters=4B2026.01 | 48.7 | |
| CODIBackbone=Qwen3-VL-8B + LoRA2026.05 | 48.6 | |
| GRPO w/o ToolsModel size=< 7B parameters, Tool-access setting=removed, Reasoning-stage setting=GRPO2026.01 | 48.5 | |
| Qwen2.5VLBackbone=Qwen2.5VL-7B-Instruct2025.12 | 48.15 | |
| Cold-start w/o ReasoningModel size=< 7B parameters, Reasoning-stage setting=removed, SFT-stage setting=Cold-start2026.01 | 47.5 | |
| Direct GRPO w/o cold-startModel size=< 7B parameters, SFT-stage setting=removed, Reasoning-stage setting=Direct GRPO2026.01 | 47 | |
| Cold-start w/o ToolsModel size=< 7B parameters, Tool-access setting=removed, SFT-stage setting=Cold-start2026.01 | 46.5 | |
| Med-R1Category=MLLMs can Think about Images2026.04 | 45.8 | |
| QWEN3-VL-8B-INSTRUCT + MED-SCOUTParameters=8B, Enhancement=Med-Scout2026.01 | 45.5 | |
| QWEN3-VL-4B-INSTRUCT + MED-SCOUTParameters=4B, Enhancement=Med-Scout2026.01 | 45.1 | |
| MedVLM-R1Category=MLLMs can Think about Images2026.04 | 44.8 | |
| Internvl3-2BModel size=< 7B parameters2026.01 | 44.5 | |
| QWEN3-VL-8B-INSTRUCTParameters=8B2026.01 | 43.9 | |
| MEDVISTAGYM (InternVL3-8B)Model size=7-13B parameters, Tool-access setting=enabled2026.01 | 43.5 | |
| BioMediX2-8BModel Category=Medical Models2026.04 | 43.5 | |
| TMA-AllComponBase Model=Gemma-3-4B2025.08 | 43.2 | |
| QWEN3-VL-4B-INSTRUCTParameters=4B2026.01 | 42.8 | |
| MEDVISTAGYM (InternVL3-2B)Model size=< 7B parameters, Tool-access setting=enabled2026.01 | 40 | |
| TMA-AllComponBase Model=MedGemma-4B2025.08 | 38.5 | |
| LLaVA-Next-13BCategory=Open-Source SOTA2026.04 | 36.6 | |
| LLaVA-v1.5-8BCategory=Open-Source SOTA2026.04 | 36.4 | |
| LLaVa1.6-7BModel Category=General Models2026.04 | 36.4 | |
| LLaVA-Next-7BCategory=Open-Source SOTA2026.04 | 35.5 | |
| Gemini-2.5-flashModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 35 | |
| LLAVA-MED-7BParameters=7B2026.01 | 32.4 | |
| VividMed2026.05 | 31.15 | |
| LLaVA-MedSize=7B2026.03 | 30.5 | |
| LLaVa-Med-7BModel Category=Medical Models2026.04 | 30.5 | |
| LVRWeights=Original2026.05 | 30.4 | |
| LLaVA-Med-7BModel size=7-13B parameters2026.01 | 27 | |
| RadFMCategory=Medical MLLMs2026.04 | 25.9 | |
| MCOUTWeights=Original2026.05 | 25.6 | |
| LLaVA-Med-7BCategory=Medical MLLMs2026.04 | 24.7 |