Medical Visual Question Answering on PMC-VQA (test)
84.6AccuracyMedCausalX
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MedCausalXfine-tuned=true, evaluation=5-fold cross validation2026.03 | 84.6 | 80.2 | 35.8 | |
| MedVLM-R1fine-tuned=true, evaluation=5-fold cross validation2026.03 | 82.4 | 75.2 | 45.8 | |
| Med-R1fine-tuned=true, evaluation=5-fold cross validation2026.03 | 81.9 | 74.5 | 47.2 | |
| MedCoTfine-tuned=true, evaluation=5-fold cross validation2026.03 | 80.3 | 71.8 | 50.5 | |
| MedRegAfine-tuned=false, evaluation=5-fold cross validation2026.03 | 79.5 | 72.8 | 49.8 | |
| MedDrfine-tuned=false, evaluation=5-fold cross validation2026.03 | 66.2 | 66.5 | 54.5 | |
| LLaVA-Medfine-tuned=false, evaluation=5-fold cross validation2026.03 | 63.8 | 55.5 | 65.2 | |
| InternVLfine-tuned=false, evaluation=5-fold cross validation2026.03 | 61.5 | 59.2 | 60.8 | |
| Qwen2.5-VLfine-tuned=false, evaluation=5-fold cross validation2026.03 | 58.2 | 56.7 | 63.5 | |
| MedEyesBackbone=Qwen2.5-VL-3B, Visual Expert=MedPLIB, Resolution=336x336, Patch size=142025.11 | 55.3 | — | — | |
| GMAI-VLCategory=Medical-Specific Models2025.11 | 52.3 | — | — | |
| HuatuoGPT-V-7BModel Category=Medical VLM2025.08 | 50 | — | — | |
| InternVL3-14BModel Category=General VLM2025.08 | 48.75 | — | — | |
| ARMed-RModel Category=Fine-tuned VLM2025.08 | 48.75 | — | — | |
| InternVL3-8BModel Category=General VLM2025.08 | 48.55 | — | — | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=GRPO2025.08 | 48.1 | — | — | |
| Qwen2.5-VL-7BModel Category=General VLM2025.08 | 47.8 | — | — | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=SFT2025.08 | 46.8 | — | — | |
| Med-R1Category=Reinforcement Learning Methods2025.11 | 45.8 | — | — | |
| RadFMfine-tuned=false, evaluation=5-fold cross validation2026.03 | 45.8 | 56.2 | 62.8 | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=Base2025.08 | 45.7 | — | — | |
| DeepEyes†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 45.2 | — | — | |
| MedVLM-R1Category=Reinforcement Learning Methods2025.11 | 44.8 | — | — | |
| GRIT†Category=Reinforcement Learning Methods, Adaptation=Using the same dataset adapted to medical domains2025.11 | 42.3 | — | — | |
| GPT-4oCategory=General Vision-Language Models2025.11 | 40.8 | — | — | |
| InternVL3-2BModel Category=General VLM2025.08 | 40.69 | — | — | |
| InternVL-2Category=General Vision-Language Models2025.11 | 38.4 | — | — | |
| Qwen2.5-VL-3BCategory=General Vision-Language Models2025.11 | 37.5 | — | — | |
| Med-Flamingofine-tuned=false, evaluation=5-fold cross validation2026.03 | 35.9 | 40.5 | 73.8 | |
| Med-FlamingoCategory=Medical-Specific Models2025.11 | 34.7 | — | — | |
| LLaVA-v1.6-13BModel Category=General VLM2025.08 | 34.3 | — | — | |
| LLaVA-v1.6-7BModel Category=General VLM2025.08 | 33.05 | — | — | |
| RadFMCategory=Medical-Specific Models2025.11 | 25.9 | — | — | |
| LLaVA-MedCategory=Medical-Specific Models2025.11 | 24.7 | — | — | |
| LLaVA-Med-7BModel Category=Medical VLM2025.08 | 23.8 | — | — | |
| MedVInTCategory=Medical-Specific Models2025.11 | 23.3 | — | — |