Visual Question Answering on PathVQA
92.9Accuracy (Closed)NVILA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| NVILASize=8B2024.12 | 92.9 | — | |
| BioMed-VITALModel size=13b, Training sample size=150K, Evaluation protocol=Supervised fine-tuning2024.06 | 92.42 | 39.89 | |
| LLaVA-MedModel size=13b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 92.39 | 38.82 | |
| Task-specific SOTA2024.12 | 91.7 | — | |
| BioMed-VITALModel size=13b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 91.41 | 39.71 | |
| LLaVA-MedModel size=7b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 91.21 | 37.95 | |
| VILA-M3Size=8B2024.12 | 91 | — | |
| M2I2Evaluation protocol=Literature-reported2024.06 | 88 | 36.3 | |
| Prefix T. Medical LMEvaluation protocol=Literature-reported2024.06 | 87 | 40 | |
| MMQEvaluation protocol=Literature-reported2024.06 | 84 | 13.4 | |
| Med-Gemini2024.12 | 83.3 | — | |
| MUMCEvaluation protocol=Literature-reported2024.06 | 65.1 | 39 | |
| LLaVAModel size=7b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 63.2 | 7.74 | |
| Shazam2025.03 | 57.5 | — | |
| UNI 22025.03 | 56.5 | — | |
| Prov-Gigapath2025.03 | 56.2 | — | |
| Virchow22025.03 | 55.9 | — | |
| H-optimus-12025.03 | 55.1 | — | |
| Phikon-v22025.03 | 53.2 | — |