Visual Question Answering on Slake
91.1Closed AccuracyM2I2
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| M2I2Evaluation protocol=Literature-reported2024.06 | 91.1 | — | — | 74.7 | |
| BioMed-VITALModel size=13b, Training sample size=150K, Evaluation protocol=Supervised fine-tuning2024.06 | 90.7 | — | — | 91.69 | |
| BiomedCLIPEvaluation protocol=Literature-reported2024.06 | 89.7 | — | — | 82.05 | |
| PMC-CLIP2023.03 | 88 | 81.9 | 84.3 | — | |
| PMC-CLIPEvaluation protocol=Literature-reported2024.06 | 88 | — | — | 81.9 | |
| PMC-CLIPPretraining Data=PMC-OA [32]2023.05 | 88 | 81.8 | — | — | |
| M3AE2023.03 | 87.82 | 80.31 | 83.25 | — | |
| M3AEEvaluation protocol=Literature-reported2024.06 | 87.82 | — | — | 80.31 | |
| M3AEPretraining Data=CC12M [9]2023.05 | 87.8 | 80.3 | — | — | |
| MedVInT-TEPretraining Data=PMC-VQA2023.05 | 87.7 | 88.2 | — | — | |
| BioMed-VITALModel size=13b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 86.54 | — | — | 87.82 | |
| MedVInT-TDPretraining Data=PMC-VQA2023.05 | 86.3 | 84.5 | — | — | |
| LLaVA-MedModel size=13b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 85.58 | — | — | 84.97 | |
| LLaVA-MedModel size=7b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 85.34 | — | — | 83.08 | |
| MedVInT-TE-SPretraining Data=None2023.05 | 85.1 | 84 | — | — | |
| MedVInT-TD-SPretraining Data=None2023.05 | 85.1 | 79.7 | — | — | |
| CPRD-BAN2023.03 | 83.4 | 79.5 | 81.1 | — | |
| M3AEPretraining Data=None2023.05 | 83.4 | 58.2 | — | — | |
| CPRD-BANPretraining Data=ROCO, MedICaT [46, 52]2023.05 | 83.4 | 78.3 | — | — | |
| PubMedCLIPEvaluation protocol=Literature-reported2024.06 | 82.5 | — | — | 78.4 | |
| Prefix T. Medical LMEvaluation protocol=Literature-reported2024.06 | 82.01 | — | — | 84.3 | |
| MUMCEvaluation protocol=Literature-reported2024.06 | 81.5 | — | — | 81.5 | |
| PMC-CLIPPretraining Data=None2023.05 | 80 | 72.7 | — | — | |
| MEVF-BAN2023.03 | 79.8 | 77.8 | 78.6 | — | |
| MEVF-BANPretraining Data=VQA-RAD [28]2023.05 | 79.8 | 77.8 | — | — | |
| CoQAHEvaluation protocol=Literature-reported2024.06 | 73.9 | — | — | 42.5 | |
| LLaVAModel size=7b, Training sample size=60K, Evaluation protocol=Supervised fine-tuning2024.06 | 63.22 | — | — | 78.18 |