Multi-image Medical Visual Question Answering on MIM-RAD (Held-out)
0.69Close Score (C)GPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o2025.05 | 0.69 | 0.3952 | |
| Med-Mantistraining_dataset=Med-MIM instruction dataset2025.05 | 0.6533 | 0.3535 | |
| Mantis-8B2025.05 | 0.5967 | 0.2808 | |
| InternVL2-8B2025.05 | 0.5633 | 0.3562 | |
| MIM-LLAVA-Medtraining_dataset=Med-MIM instruction dataset2025.05 | 0.5333 | 0.3508 | |
| LLaVA-Med-7B2025.05 | 0.4867 | 0.3486 | |
| deepseek-VL-7B2025.05 | 0.17 | 0.3158 | |
| Mantis†setup=without composed instruction dataset2025.05 | 0.0566 | 0.0727 | |
| LLaVA-Med†setup=without composed instruction dataset2025.05 | 0.0466 | 0.0022 | |
| Med-Flamingo-9B2025.05 | 0.04 | 0.2869 |