Multi-image Medical Visual Question Answering on MIM-ODIR (Held-out)
45.67VQA Close Accuracy (C)GPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o2025.05 | 45.67 | 34.6 | |
| MIM-LLAVA-Medtraining_dataset=Med-MIM instruction dataset2025.05 | 32.33 | 55.99 | |
| InternVL2-8B2025.05 | 31.67 | 41.01 | |
| Med-Mantistraining_dataset=Med-MIM instruction dataset2025.05 | 30 | 41.32 | |
| LLaVA-Med-7B2025.05 | 29.67 | 53.89 | |
| Mantis-8B2025.05 | 23 | 21.2 | |
| deepseek-VL-7B2025.05 | 21.67 | 44.49 | |
| Med-Flamingo-9B2025.05 | 16 | 30.83 | |
| Mantis†setup=without composed instruction dataset2025.05 | 7 | 20.12 | |
| LLaVA-Med†setup=without composed instruction dataset2025.05 | 2.66 | 2.1 |