Medical Visual Question Answering on Average (OmniMedVQA, PMC-VQA, MedXpertQA) cross-dataset (test)
0.155ECEOurs
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OursModel Architecture=MedGemma-4B-IT, Training Pathway=OmniMedVQA-trained2026.06 | 0.155 | 0.204 | 0.693 | 0.501 | |
| SteerConfModel Architecture=MedGemma-4B-IT, Training Pathway=OmniMedVQA-trained2026.06 | 0.28 | 0.327 | 0.603 | 0.465 | |
| ConfTunerModel Architecture=MedGemma-4B-IT, Training Pathway=OmniMedVQA-trained2026.06 | 0.311 | 0.339 | 0.599 | 0.449 | |
| Base ModelModel Architecture=MedGemma-4B-IT, Training Pathway=OmniMedVQA-trained2026.06 | 0.416 | 0.427 | 0.549 | 0.466 | |
| Top-K SamplingModel Architecture=MedGemma-4B-IT, Training Pathway=OmniMedVQA-trained2026.06 | 0.553 | 0.525 | 0.52 | 0.321 | |
| Ours (PMC-trained)Model=MedGemma, Training Dataset=PMC-VQA2026.06 | 12.3 | 22.6 | 62.4 | 46.1 | |
| SteerConfModel=MedGemma2026.06 | 28 | 32.7 | 60.3 | 46.5 | |
| Base ModelModel=MedGemma2026.06 | 41.6 | 42.7 | 54.9 | 46.6 | |
| ConfTunerModel=MedGemma2026.06 | 43.7 | 44.5 | 56.3 | 45.9 | |
| Top-K SamplingModel=MedGemma2026.06 | 55.3 | 52.5 | 52 | 32.1 |