Medical Visual Question Answering on PathVQA (Open and Closed Metrics)
72.61Overall AccuracyUpper Bound
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Upper BoundNoise Type=Clean2025.03 | 72.61 | 52.64 | 86.45 | — | |
| GPT-4oDecoding=Greedy2025.08 | 72.43 | — | — | — | |
| Cold-start w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, SFT-stage setting=Cold-start2026.01 | 72 | — | — | — | |
| Gemini-2.5-proModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 71 | — | — | — | |
| GRPO w/o ToolsModel size=7-13B parameters, Tool-access setting=removed, Reasoning-stage setting=GRPO2026.01 | 70.5 | — | — | — | |
| IBISAgent2026.01 | 69.2 | — | — | — | |
| MEDVISTA-R1Model size=7-13B parameters2026.01 | 69 | — | — | — | |
| MEDVISTA-R1Tool=true2026.01 | 69 | — | — | — | |
| Chiron2026.01 | 68.9 | — | — | — | |
| MedVLThinker-32B RLDecoding=Greedy2025.08 | 68.82 | — | — | — | |
| Chiron-01-8BTool=false, Category=Medical MLLMs with CoT Reasoning2026.01 | 68.8 | — | — | — | |
| Lingshu2026.01 | 68.7 | — | — | — | |
| GPT-5-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 68.65 | — | — | — | |
| Lingshu-7BTool=false, Category=Medical MLLMs with CoT Reasoning2026.01 | 68.4 | — | — | — | |
| InternVL3-8BModel size=7-13B parameters2026.01 | 68.1 | — | — | — | |
| Direct GRPO w/o cold-startModel size=7-13B parameters, SFT-stage setting=removed, Reasoning-stage setting=Direct GRPO2026.01 | 68 | — | — | — | |
| Qwen2.5-VL-32B-InstructDecoding=Greedy2025.08 | 67.98 | — | — | — | |
| MedVLThinker-7B RLDecoding=Greedy2025.08 | 66.83 | — | — | — | |
| HuatuoGPT-Vision-34BDecoding=Greedy2025.08 | 66.72 | — | — | — | |
| DiNNoise Type=10%-Semantic Noise2025.03 | 66.67 | 40.54 | 80.28 | — | |
| VILA-M3-40BTool=true, Category=Multimodal medical agents2026.01 | 66.4 | — | — | — | |
| DCI2026.04 | 66.4 | 40.6 | 92 | — | |
| GPT-o4-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 66 | — | — | — | |
| Cold-start w/o ReasoningModel size=7-13B parameters, Reasoning-stage setting=removed, SFT-stage setting=Cold-start2026.01 | 66 | — | — | — | |
| Direct GRPO w/o cold-startModel size=< 7B parameters, SFT-stage setting=removed, Reasoning-stage setting=Direct GRPO2026.01 | 66 | — | — | — | |
| Gemme 3 27BDecoding=Greedy2025.08 | 65.7 | — | — | — | |
| Qwen2.5-VL-7B-InstructDecoding=Greedy2025.08 | 65.39 | — | — | — | |
| MUMC2026.04 | 65.1 | 39 | 90.4 | — | |
| GPT-5Model size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 64.4 | — | — | — | |
| Q2ATransformerNoise Type=10%-Semantic Noise2025.03 | 64.18 | 38.85 | 78.33 | — | |
| CLIP-ViT2026.04 | 63.6 | 40 | 87 | — | |
| HuatuoGPT-Vision-7BDecoding=Greedy2025.08 | 63.53 | — | — | — | |
| GPT-4o-miniDecoding=Greedy2025.08 | 63.33 | — | — | — | |
| Claude-4.5-sonnetModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 63 | — | — | — | |
| GRPO w/o ToolsModel size=< 7B parameters, Tool-access setting=removed, Reasoning-stage setting=GRPO2026.01 | 63 | — | — | — | |
| Qwen3vl-8BModel size=7-13B parameters2026.01 | 62.9 | — | — | — | |
| Q2ATransformerNoise Type=Clean2025.03 | 62.61 | 38.42 | 76.85 | — | |
| Qwen2.5vl-7BModel size=7-13B parameters2026.01 | 62.4 | — | — | — | |
| BaselineNoise Type=Clean2025.03 | 62.32 | 37.83 | 75.81 | — | |
| MedVLThinker-3B RLDecoding=Greedy2025.08 | 62.28 | — | — | — | |
| M2I22026.04 | 62.2 | 36.3 | 88 | — | |
| MedGemma 27BDecoding=Greedy2025.08 | 62.09 | — | — | — | |
| BaselineNoise Type=10%-Semantic Noise2025.03 | 62.06 | 36.54 | 75.35 | — | |
| Qwen2.5-VL-3B-InstructDecoding=Greedy2025.08 | 61.96 | — | — | — | |
| MMBERTNoise Type=10%-Semantic Noise2025.03 | 61.02 | 32.88 | 75.28 | — | |
| Cold-start w/o ToolsModel size=< 7B parameters, Tool-access setting=removed, SFT-stage setting=Cold-start2026.01 | 61 | — | — | — | |
| DiNNoise Type=20%-Random Noise2025.03 | 60.84 | 37.85 | 76.16 | — | |
| Claude-4.5-haikuModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 60 | — | — | — | |
| AURATool=true, Category=Multimodal medical agents2026.01 | 59.8 | — | — | — | |
| SNLCNoise Type=20%-Random Noise2025.03 | 59.71 | 36.74 | 75.03 | — | |
| MedGemma 4BDecoding=Greedy2025.08 | 59.64 | — | — | — | |
| MMedAgent-7BTool=true, Category=Multimodal medical agents2026.01 | 59.47 | — | — | — | |
| Gemme 3 4BDecoding=Greedy2025.08 | 59.24 | — | — | — | |
| SimTNoise Type=20%-Random Noise2025.03 | 59.01 | 35.17 | 74.91 | — | |
| LLaVA-Med-7BModel size=7-13B parameters2026.01 | 59 | — | — | — | |
| CoDisNoise Type=20%-Random Noise2025.03 | 58.96 | 35.66 | 74.5 | — | |
| MMBERTNoise Type=Clean2025.03 | 58.73 | 34.18 | 73.21 | — | |
| Q2ATransformerNoise Type=20%-Random Noise2025.03 | 58.64 | 34.5 | 74.74 | — | |
| MedAgent-ProTool=true, Category=Multimodal medical agents2026.01 | 58.5 | — | — | — | |
| MMedAgentRL-7BTool=false, Category=Multimodal medical agents2026.01 | 58.5 | — | — | — | |
| DiNNoise Type=20%-Semantic Noise2025.03 | 58.14 | 32.47 | 72.56 | — | |
| GPT-4o2026.01 | 58.1 | — | — | — | |
| BaselineNoise Type=20%-Random Noise2025.03 | 57.03 | 32.29 | 73.53 | — | |
| Cold-start w/o ReasoningModel size=< 7B parameters, Reasoning-stage setting=removed, SFT-stage setting=Cold-start2026.01 | 57 | — | — | — | |
| SNLCNoise Type=20%-Semantic Noise2025.03 | 56.72 | 30.88 | 71.23 | — | |
| CoDisNoise Type=20%-Semantic Noise2025.03 | 56.55 | 30.36 | 71.18 | — | |
| Llava Med v1.5 Mistral 7BDecoding=Greedy2025.08 | 56.52 | — | — | — | |
| MEDVISTAGYM (Qwen3vl-8B)Model size=7-13B parameters, Tool-access setting=enabled2026.01 | 56.4 | — | — | — | |
| Q2ATransformerNoise Type=20%-Semantic Noise2025.03 | 56.28 | 30.16 | 71.11 | — | |
| SimTNoise Type=20%-Semantic Noise2025.03 | 55.93 | 29.38 | 70.69 | — | |
| Lingshu-7BType=Comprehension only, # Params=7B, # Data=7.1M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 55.5 | — | 85 | — | |
| MMBERTNoise Type=20%-Random Noise2025.03 | 55.41 | 30.45 | 72.06 | — | |
| BaselineNoise Type=20%-Semantic Noise2025.03 | 55.28 | 29.06 | 70.01 | — | |
| MEDVISTA-R1Model size=< 7B parameters2026.01 | 55 | — | — | — | |
| Mini-o3-7B-vlTool=true, Category=MLLMs can think with image2026.01 | 53.4 | — | — | — | |
| DeepEyes-7BTool=true, Category=MLLMs can think with image2026.01 | 52.9 | — | — | — | |
| PixelReasoner-RL-vl-7BTool=true, Category=MLLMs can think with image2026.01 | 52.6 | — | — | — | |
| MMBERTNoise Type=20%-Semantic Noise2025.03 | 52.41 | 27.13 | 68.35 | — | |
| HuatuoGPT-Vision-34B2026.01 | 51.3 | — | — | — | |
| HuatuoGPT-Vision-34BTool=false, Category=Medical MLLMs2026.01 | 50.7 | — | — | — | |
| AMAM2026.04 | 50.4 | 18.2 | 84.4 | — | |
| Qwen2.5vl-32BTool=false, Category=Opensource SOTA2026.01 | 47.4 | — | — | — | |
| MEDVISTAGYM (InternVL3-8B)Model size=7-13B parameters, Tool-access setting=enabled2026.01 | 46 | — | — | — | |
| MEVF-BAN2026.04 | 44.8 | 8.1 | 81.4 | — | |
| LlaVa-med-7BTool=false, Category=Medical MLLMs2026.01 | 44.6 | — | — | — | |
| HealthGPT-L14Type=Comprehension only, # Params=14B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 44.4 | — | 85.9 | — | |
| LLaVA-Med-7B2026.01 | 44.2 | — | — | — | |
| MEDSIGHTType=Unified, # Params=8B, # Data=72K, Evaluation Protocol=Zero-shot inference2026.06 | 42.6 | — | 66.3 | — | |
| Internvl3-2BModel size=< 7B parameters2026.01 | 41.8 | — | — | — | |
| HuatuoGPT-VisionType=Comprehension only, # Params=7B, # Data=647K, Evaluation Protocol=Zero-shot inference2026.06 | 41.3 | — | 63.6 | — | |
| Llava-Next-13BTool=false, Category=Opensource SOTA2026.01 | 39.8 | — | — | — | |
| HealthGPT-M3Type=Comprehension only, # Params=3.8B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 39.7 | — | 78.7 | — | |
| InternVL3.5Type=Comprehension only, # Params=8B, # Data=16.3M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 38.6 | — | 67 | — | |
| MedVLM-R1-2BTool=false, Category=Medical MLLMs with CoT Reasoning2026.01 | 38.3 | — | — | — | |
| SMR-AgentsTool=true, Category=Multimodal medical agents2026.01 | 38.2 | — | — | — | |
| Qwen3-VLType=Comprehension only, # Params=8B, # Data=-, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 36.9 | — | 66.9 | — | |
| OMG-LLaVAType=Unified, # Params=7B, # Data=1.2M, Evaluation Protocol=Zero-shot inference2026.06 | 36.2 | — | 65.2 | — | |
| Gemini-2.5-flashModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 36.08 | — | — | — | |
| LLaVA-MedType=Comprehension only, # Params=7B, # Data=60K, Evaluation Protocol=Zero-shot inference2026.06 | 35.7 | — | 62.3 | — | |
| Llama-3.2Type=Comprehension only, # Params=11B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 33.6 | — | 62.8 | — |