Medical Visual Question Answering on PathVQA (test)
78.2AccuracyMeissa
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| MeissaFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 78.2 | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 74.3 | — | — | — | — | — | — | |
| Ophiuchus-7BFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 74.3 | — | — | — | — | — | — | |
| MDAgentsCategory=Multi-Agents2026.03 | 73.97 | — | — | — | — | — | — | |
| Qwen3-VL-235BCategory=Single LLM2026.03 | 73.77 | — | — | — | — | — | — | |
| Intern-S1Category=Single LLM2026.03 | 73.74 | — | — | — | — | — | — | |
| MedChain-AgentsCategory=Multi-Agents2026.03 | 73.56 | — | — | — | — | — | — | |
| AutoGenCategory=Multi-Agents2026.03 | 73.44 | — | — | — | — | — | — | |
| MedCausalXfine-tuned=true, evaluation=5-fold cross validation2026.03 | 73.2 | — | — | 75.8 | 38.9 | — | — | |
| Qwen3-VL-4BFramework=Framework II, Inference Mode=Direct Inference, Training Protocol=SFT2026.03 | 73 | — | — | — | — | — | — | |
| ClinicalAgentsCategory=Multi-Agents2026.03 | 72.78 | — | — | — | — | — | — | |
| ReActCategory=Single Agent2026.03 | 71.74 | — | — | — | — | — | — | |
| ReConcileCategory=Multi-Agents2026.03 | 71.24 | — | — | — | — | — | — | |
| RAGCategory=Single Agent2026.03 | 71.12 | — | — | — | — | — | — | |
| MedVLM-R1fine-tuned=true, evaluation=5-fold cross validation2026.03 | 70.8 | — | — | 70.5 | 49.8 | — | — | |
| Few-shot + CoTCategory=Single Agent2026.03 | 70.17 | — | — | — | — | — | — | |
| Med-R1fine-tuned=true, evaluation=5-fold cross validation2026.03 | 69.5 | — | — | 69.2 | 51.5 | — | — | |
| ColaCareCategory=Multi-Agents2026.03 | 69.01 | — | — | — | — | — | — | |
| Llama-4-Maverick-17BCategory=Single LLM2026.03 | 68.8 | — | — | — | — | — | — | |
| MedRegAfine-tuned=false, evaluation=5-fold cross validation2026.03 | 68.5 | — | — | 68.5 | 52.5 | — | — | |
| GPT-5.2Category=Single LLM2026.03 | 68.35 | — | — | — | — | — | — | |
| MedAgentsCategory=Multi-Agents2026.03 | 68.29 | — | — | — | — | — | — | |
| MeissaFramework=Framework III, Inference Mode=Agent-based, Agent System=MDAgents2026.03 | 67.9 | — | — | — | — | — | — | |
| o3Framework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 67.5 | — | — | — | — | — | — | |
| Gemini-2.5-proFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 67.1 | — | — | — | — | — | — | |
| Qwen3-VL-4BFramework=Framework III, Inference Mode=Direct Inference, Training Protocol=SFT2026.03 | 66.4 | — | — | — | — | — | — | |
| MedCoTfine-tuned=true, evaluation=5-fold cross validation2026.03 | 66.2 | — | — | 66.8 | 54.8 | — | — | |
| Qwen3-VL-4BFramework=Framework III, Inference Mode=Agent-based, Agent System=MDAgents2026.03 | 65.5 | — | — | — | — | — | — | |
| Qwen3-VL-4BFramework=Framework II, Inference Mode=Agent-based, Agent System=Ophiuchus2026.03 | 65.3 | — | — | — | — | — | — | |
| GPT-4VFramework=Framework III, Inference Mode=Agent-based, Agent System=MDAgents2026.03 | 65.3 | — | — | — | — | — | — | |
| MedEyesBackbone=Qwen2.5-VL-3B, Visual Expert=MedPLIB, Resolution=336x336, Patch size=142025.11 | 64.8 | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework II, Inference Mode=Direct Inference2026.03 | 64.3 | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework III, Inference Mode=Direct Inference2026.03 | 64.3 | — | — | — | — | — | — | |
| MedDrfine-tuned=false, evaluation=5-fold cross validation2026.03 | 62.8 | — | — | 64.8 | 56.8 | — | — | |
| GPT-5Framework=Framework II, Inference Mode=Direct Inference2026.03 | 60 | — | — | — | — | — | — | |
| GPT-4oCategory=General Vision-Language Models2025.11 | 59.2 | — | — | — | — | — | — | |
| LLaVA-Medfine-tuned=false, evaluation=5-fold cross validation2026.03 | 58.2 | — | — | 53.2 | 63.8 | — | — | |
| GPT-4VFramework=Framework III, Inference Mode=Direct Inference2026.03 | 57.9 | — | — | — | — | — | — | |
| LLaVA-MedCategory=Medical-Specific Models2025.11 | 56.8 | — | — | — | — | — | — | |
| Gemini-3-flashFramework=Framework III, Inference Mode=Agent-based, Agent System=MDAgents2026.03 | 56.3 | — | — | — | — | — | — | |
| Qwen2.5-VL-3BCategory=General Vision-Language Models2025.11 | 55.2 | — | — | — | — | — | — | |
| MedVLM-R1Category=Reinforcement Learning Methods2025.11 | 55.2 | — | — | — | — | — | — | |
| MedVInTCategory=Medical-Specific Models2025.11 | 54.7 | — | — | — | — | — | — | |
| MedGemma-4BCategory=Single LLM2026.03 | 54.55 | — | — | — | — | — | — | |
| Med-R1Category=Reinforcement Learning Methods2025.11 | 53.3 | — | — | — | — | — | — | |
| InternVLfine-tuned=false, evaluation=5-fold cross validation2026.03 | 52.7 | — | — | 55.3 | 65.3 | — | — | |
| Qwen2.5-VLfine-tuned=false, evaluation=5-fold cross validation2026.03 | 48.3 | — | — | 52.8 | 68.2 | — | — | |
| Med-Flamingofine-tuned=false, evaluation=5-fold cross validation2026.03 | 47.9 | — | — | 45.6 | 68.5 | — | — | |
| GMAI-VLCategory=Medical-Specific Models2025.11 | 47.2 | — | — | — | — | — | — | |
| InternVL-2Category=General Vision-Language Models2025.11 | 45.8 | — | — | — | — | — | — | |
| GRIT†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 43.5 | — | — | — | — | — | — | |
| DeepEyes†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 42.3 | — | — | — | — | — | — | |
| Med-FlamingoCategory=Medical-Specific Models2025.11 | 40.7 | — | — | — | — | — | — | |
| RadFMfine-tuned=false, evaluation=5-fold cross validation2026.03 | 40.2 | — | — | 50.8 | 66.5 | — | — | |
| RadFMCategory=Medical-Specific Models2025.11 | 38.7 | — | — | — | — | — | — | |
| CLIP-ViT2025.05 | — | 40 | 87 | — | — | 63.6 | — | |
| DeRS-LMMoE Model=Med-MoE-StableLM, Added Params=5.63M2025.03 | — | 33.9 | 91.4 | — | — | — | — | |
| DeRS-LMMoE Model=Med-MoE-Phi, Added Params=9.18M2025.03 | — | 35.6 | 91.9 | — | — | — | — | |
| DeRS-SMMoE Model=Med-MoE-StableLM, Added Params=2.17M2025.03 | — | 33.6 | 90.9 | — | — | — | — | |
| DeRS-SMMoE Model=Med-MoE-Phi, Added Params=5.18M2025.03 | — | 35 | 91.6 | — | — | — | — | |
| Gemini 2.5 ProProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | — | — | 41.4 | — | — | — | 14.2 | |
| GLM-4.5VProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | — | — | 64.6 | — | — | — | 23.8 | |
| GLM-4VProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 74.7 | — | — | — | 24.2 | |
| GPT-5Protocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | — | — | 73.1 | — | — | — | 28.5 | |
| InternVL-30BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 71.9 | — | — | — | 21.8 | |
| InternVL-38BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 70.5 | — | — | — | 19.6 | |
| InternVL-6BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 66.2 | — | — | — | 17 | |
| LLaVAProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 13.51 | — | — | — | 6.26 | |
| LLaVAProtocol=Supervised fine-tuned2026.04 | — | — | 63.2 | — | — | — | 7.74 | |
| LLaVAParams=7B2026.05 | — | — | 90.8 | — | — | — | 37.1 | |
| LLaVA-Med2025.05 | — | 38.9 | 91.7 | — | — | 65.3 | — | |
| LLaVA-MedProtocol=Supervised fine-tuned2026.04 | — | — | 91.21 | — | — | — | 37.95 | |
| LLaVA-MedParams=7B2026.05 | — | — | 91.2 | — | — | — | 38.5 | |
| LLaVA-TriParams=8B2026.05 | — | — | 91.8 | — | — | — | 37.8 | |
| M2I22025.05 | — | 36.3 | 88 | — | — | 62.2 | — | |
| MEVF-BAN2025.05 | — | 8.1 | 81.4 | — | — | 44.8 | — | |
| MMADAProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 40.8 | — | — | — | 1.2 | |
| o3Protocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | — | — | 74.1 | — | — | — | 29.4 | |
| OursModel Scale=7B2025.05 | — | 40.4 | 87.9 | — | — | 64.2 | — | |
| OursModel Scale=13B2025.05 | — | 41.4 | 91.5 | — | — | 66.5 | — | |
| parameter-efficient multi-level prompt frameworkParams=104M2026.05 | — | — | 91.3 | — | — | — | 37.2 | |
| PeFoMedParams=33M, Configuration=LoRA2026.05 | — | — | 88.4 | — | — | — | 29.5 | |
| PREFIX T. MEDICAL LMModel Type=Representative non-VLM2026.04 | — | — | 87 | — | — | — | 40 | |
| Q2ATRANSFORMERModel Type=Representative non-VLM2026.04 | — | — | 88.85 | — | — | — | 54.85 | |
| Qwen 2.5 VLProtocol=Zero-shot, Model Type=frontier and grounding-aware VLM2026.04 | — | — | 71.3 | — | — | — | 26.7 | |
| Qwen 2.5 VL (Ours, SFT)Protocol=Supervised fine-tuned2026.04 | — | — | 90.3 | — | — | — | 40.7 | |
| Qwen2.5-VL-7BProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 10.3 | — | — | — | 21.5 | |
| THYMEProtocol=Zero-shot, Model Type=open-source VLM2026.04 | — | — | 63.9 | — | — | — | 9.9 | |
| VanillaMoE Model=Med-MoE-StableLM, Added Params=1.66B2025.03 | — | 33.4 | 91.4 | — | — | — | — | |
| VanillaMoE Model=Med-MoE-Phi, Added Params=3.36B2025.03 | — | 35.1 | 91.5 | — | — | — | — | |
| VL ENCODER–DECODERModel Type=Representative non-VLM2026.04 | — | — | 84.63 | — | — | — | 58.29 |