Medical Visual Question Answering on VQA-RAD
80.4AccuracyRobust-MMR
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Robust-MMREvaluation Setting=Standard2026.02 | 80.4 | — | — | — | — | — | |
| OctoMed-7BParameters=7B2026.03 | 79 | — | — | — | — | — | |
| Ours (7B)Parameters=7B2026.03 | 79 | — | — | — | — | — | |
| MedGemma 4BDecoding=Greedy2025.08 | 78.55 | — | — | — | — | — | |
| MedGemma 4BModel Type=Medical LLM2025.10 | 78.55 | — | — | — | — | — | |
| KG-CMIreported_format=Mean ± Standard Deviation2026.04 | 78.21 | 68.18 | 84.11 | — | — | — | |
| CPRDEvaluation Setting=Standard2026.02 | 78.2 | — | — | — | — | — | |
| LaPAreported_format=Mean ± Standard Deviation2026.04 | 78.15 | 67.87 | 84.92 | — | — | — | |
| MedVLSynther-7BModel Type=Medical LLM2025.10 | 77.57 | — | — | — | — | — | |
| Ours-7BModel Category=Medical Models2026.04 | 77.16 | — | — | — | — | — | |
| MEVFEvaluation Setting=Standard2026.02 | 77 | — | — | — | — | — | |
| MedVLThinker-32B RLDecoding=Greedy2025.08 | 76.96 | — | — | — | — | — | |
| M3AEreported_format=Mean ± Standard Deviation2026.04 | 76.72 | 63.5 | 85.42 | — | — | — | |
| VG-CALFreported_format=Mean ± Standard Deviation2026.04 | 76.16 | 67.47 | 85.28 | — | — | — | |
| BANEvaluation Setting=Standard2026.02 | 76 | — | — | — | — | — | |
| MUMCreported_format=Mean ± Standard Deviation2026.04 | 75.61 | 62.57 | 84.19 | — | — | — | |
| Robust-MMREvaluation Setting=Perturbed2026.02 | 75.6 | — | — | — | — | — | |
| Qwen2.5-VL-32B-InstructDecoding=Greedy2025.08 | 75.12 | — | — | — | — | — | |
| CCIS-MVQA*reported_format=Point estimate2026.04 | 75.06 | 68.78 | 79.24 | — | — | — | |
| MedVRModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain, Backbone=Qwen2.5-VL-7B2026.04 | 74.4 | — | — | — | — | — | |
| HuatuoGPT-Vision-34BDecoding=Greedy2025.08 | 74.26 | — | — | — | — | — | |
| ToolTreeBackbone Model=GPT-4o2026.03 | 74.12 | — | — | — | — | — | |
| Qwen2.5-VL-7B-InstructParameters=7B2026.03 | 74 | — | — | — | — | — | |
| QoQ-Med-7BParameters=7B2026.03 | 74 | — | — | — | — | — | |
| Qwen3vl-8BModel size=7-13B parameters2026.01 | 73.71 | — | — | — | — | — | |
| MedVLSynther-3BModel Type=Medical LLM2025.10 | 73.53 | — | — | — | — | — | |
| IBISAgent2026.01 | 73.4 | — | — | — | — | — | |
| GPT-5-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 73.31 | — | — | — | — | — | |
| MPRreported_format=Mean ± Standard Deviation2026.04 | 73.3 | 59.3 | 79.1 | — | — | — | |
| UnICLAM*reported_format=Point estimate2026.04 | 73.2 | 59.8 | 82.6 | — | — | — | |
| InternVL3-8BModel size=7-13B parameters2026.01 | 72.91 | — | — | — | — | — | |
| Chiron2026.01 | 72.7 | — | — | — | — | — | |
| MedGemma 27BDecoding=Greedy2025.08 | 72.67 | — | — | — | — | — | |
| MedGemma 27BModel Type=Medical LLM2025.10 | 72.67 | — | — | — | — | — | |
| Direct GRPO w/o cold-startModel size=7-13B parameters, SFT-stage setting=removed, Reasoning-stage setting=Direct GRPO2026.01 | 72.64 | — | — | — | — | — | |
| ViToSBackbone=Lingshu-32B2026.06 | 72.51 | — | — | — | — | — | |
| Chiron-01-8BTool=false2026.01 | 72.5 | — | — | — | — | — | |
| MedGemma-4BModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain2026.04 | 72.5 | — | — | — | — | — | |
| Qwen2.5-VL-32BModel Category=General-Purpose VLMs, Evaluation Protocol=In-domain2026.04 | 71.8 | — | — | — | — | — | |
| Qwen2.5vl-32BTool=false2026.01 | 71.7 | — | — | — | — | — | |
| Qwen2.5-VL-32BCategory=Open-Source SOTA2026.04 | 71.7 | — | — | — | — | — | |
| Gemini-2.5-proModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 71.43 | — | — | — | — | — | |
| Lingshu-32BParams=32B, Medical=yes2025.10 | 71.4 | — | — | — | — | — | |
| ViToSBackbone=Lingshu-7B2026.06 | 71.31 | — | — | — | — | — | |
| ViToS2026.06 | 71.3 | — | — | — | — | — | |
| MedVLThinker-3B RLDecoding=Greedy2025.08 | 71.08 | — | — | — | — | — | |
| MedVLThinker-3BModel Type=Medical LLM2025.10 | 71.08 | — | — | — | — | — | |
| LINGSHU-7B + MED-SCOUTParameters=7B, Enhancement=Med-Scout2026.01 | 71 | — | — | — | — | — | |
| CPRDEvaluation Setting=Perturbed2026.02 | 71 | — | — | — | — | — | |
| MEDGEMMA-4B-ITParameters=4B2026.01 | 70.8 | — | — | — | — | — | |
| GRPO w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, Reasoning-stage setting=GRPO2026.01 | 70.75 | — | — | — | — | — | |
| Cold-start w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, SFT-stage setting=Cold-start2026.01 | 70.75 | — | — | — | — | — | |
| MEDVISTA-R1Model size=7-13B parameters2026.01 | 70.75 | — | — | — | — | — | |
| MEDVISTA-R1Tool=true2026.01 | 70.75 | — | — | — | — | — | |
| ViToSBackbone=HuatuoGPT-Vision-7B2026.06 | 70.52 | — | — | — | — | — | |
| BaselineBackbone=Lingshu-32B2026.06 | 70.52 | — | — | — | — | — | |
| VQAMix*reported_format=Point estimate2026.04 | 70.5 | 56.9 | 79.5 | — | — | — | |
| GPT-4oDecoding=Greedy2025.08 | 70.22 | — | — | — | — | — | |
| GEMINI-3-FLASHModel Type=Proprietary2026.01 | 70.2 | — | — | — | — | — | |
| HUATUOGPT-VISION-7B + MED-SCOUTParameters=7B, Enhancement=Med-Scout2026.01 | 70.1 | — | — | — | — | — | |
| ViTAR2026.06 | 70.1 | — | — | — | — | — | |
| MEVFEvaluation Setting=Perturbed2026.02 | 69.4 | — | — | — | — | — | |
| M2I2reported_format=Mean ± Standard Deviation2026.04 | 69.16 | 48.78 | 82.71 | — | — | — | |
| LINGSHU-7BParameters=7B2026.01 | 68.9 | — | — | — | — | — | |
| Qwen2.5-VL-7B-InstructDecoding=Greedy2025.08 | 68.75 | — | — | — | — | — | |
| Qwen2.5-VL-7B-InstructModel Type=General LLM2025.10 | 68.75 | — | — | — | — | — | |
| CARE-Coord-BSource Type=Open-source2026.03 | 68.29 | — | — | — | — | — | |
| BANEvaluation Setting=Perturbed2026.02 | 68.2 | — | — | — | — | — | |
| LingshuSize=7B2026.03 | 67.9 | — | — | — | — | — | |
| Lingshu-7BModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain2026.04 | 67.9 | — | — | — | — | — | |
| Lingshu-7BModel Category=Medical Models2026.04 | 67.9 | — | — | — | — | — | |
| VScanBackbone=HuatuoGPT-Vision-7B2026.06 | 67.73 | — | — | — | — | — | |
| Rform + Rexact + RDTWReward functions=Rform + Rexact + RDTW2026.04 | 67.6 | — | — | — | — | — | |
| MedGemma-4BParams=4B, Medical=yes2025.10 | 67.6 | — | — | — | — | — | |
| Claude Sonnet 4Params=Closed, Medical=no2025.10 | 67.6 | — | — | — | — | — | |
| GPT-5Model size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 67.53 | — | — | — | — | — | |
| OriginalBackbone Model=Lingshu-7B2025.12 | 67.41 | — | — | — | — | — | |
| Gemini-2.5-FlashParams=Closed, Medical=no2025.10 | 67.41 | — | — | — | — | — | |
| VisionZipBackbone=HuatuoGPT-Vision-7B2026.06 | 67.33 | — | — | — | — | — | |
| HUATUOGPT-VISION-7BParameters=7B2026.01 | 67 | — | — | — | — | — | |
| HuatuoGPT-V-7BModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain2026.04 | 67 | — | — | — | — | — | |
| HuatuoGPT-V-7BModel Category=Medical Models2026.04 | 67 | — | — | — | — | — | |
| VScanBackbone=Lingshu-32B2026.06 | 66.93 | — | — | — | — | — | |
| MMQ*reported_format=Point estimate2026.04 | 66.92 | 52 | 76.71 | — | — | — | |
| GPT-4o-miniDecoding=Greedy2025.08 | 66.91 | — | — | — | — | — | |
| BaselineBackbone=HuatuoGPT-Vision-7B2026.06 | 66.53 | — | — | — | — | — | |
| PDropBackbone=HuatuoGPT-Vision-7B2026.06 | 66.53 | — | — | — | — | — | |
| Rform + RexactReward functions=Rform + Rexact2026.04 | 66.5 | — | — | — | — | — | |
| OctoToolsBackbone Model=GPT-4o2026.03 | 66.42 | — | — | — | — | — | |
| GPT-5Model Type=Proprietary2026.01 | 66.4 | — | — | — | — | — | |
| Lingshu-7BTool=false2026.01 | 66.4 | — | — | — | — | — | |
| PSIBackbone Model=Lingshu-7B2025.12 | 66.3 | — | — | — | — | — | |
| InternVL3-14BModel Category=General-Purpose VLMs, Evaluation Protocol=In-domain2026.04 | 66.3 | — | — | — | — | — | |
| GMAI-VLParams=7B, Medical=yes2025.10 | 66.3 | — | — | — | — | — | |
| MMedAgentRL-7BTool=false2026.01 | 66.1 | — | — | — | — | — | |
| Lingshu2026.01 | 66.1 | — | — | — | — | — | |
| MMedAgent-RL-7BCategory=Multimodal Medical Agents2026.04 | 66.1 | — | — | — | — | — | |
| Cold-start w/o ReasoningModel size=7-13B parameters, Reasoning-stage setting=removed, SFT-stage setting=Cold-start2026.01 | 66.04 | — | — | — | — | — | |
| PixelReasoner-RL-vl-7BTool=true2026.01 | 66 | — | — | — | — | — | |
| OpenAI-o3Category=Close-Source SOTA2026.04 | 66 | — | — | — | — | — |