Multi-modal Multi-span Medical Question Answering on M3QuestionIng
94.34AccuracyM3QAFrame
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| M3QAFrameTraining Strategy=Trained on M3QuestionIng2026.05 | 94.34 | 91.32 | 94.34 | |
| GPT-4oEvaluation Protocol=Average of zero-shot and few-shot2026.05 | 81.2 | 78.5 | 79.5 | |
| Qwen2.5-VLTraining Strategy=finetuned2026.05 | 65.38 | 5.07 | 65.38 | |
| MMEmbedEvaluation Protocol=Average of zero-shot and few-shot2026.05 | 52.15 | 66.67 | 34.59 | |
| LLaVAEvaluation Protocol=Average of zero-shot and few-shot2026.05 | 40.56 | 56.32 | 40.87 | |
| MedGemmaTraining Strategy=finetuned2026.05 | 39.16 | 55.05 | 39.16 | |
| VLM2VecEvaluation Protocol=Average of zero-shot and few-shot2026.05 | 28.05 | 43.51 | 27.84 | |
| Uni-MedCLIPEvaluation Protocol=Average of zero-shot and few-shot2026.05 | 12.74 | 22.6 | 11.78 |