Medical Visual Question Answering on MMMU Med
62.7Average ScoreM3LLM-8B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| M3LLM-8BParameters=8B2025.11 | 62.7 | 63.3 | 70 | 53.3 | 53.3 | 73.3 | |
| Lingshu-7BType=Comprehension only, # Params=7B, # Data=7.1M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 60.7 | — | — | — | — | — | |
| InternVL3-8BParameters=8B2025.11 | 57.3 | 53.3 | 66.7 | 43.3 | 60 | 63.3 | |
| QWen2.5-VL-7BParameters=7B2025.11 | 54.7 | 56.7 | 66.7 | 36.7 | 56.7 | 56.7 | |
| Lingshu-7BParameters=7B2025.11 | 54 | 56.7 | 53.3 | 60 | 46.7 | 53.3 | |
| MEDSIGHTType=Unified, # Params=8B, # Data=72K, Evaluation Protocol=Zero-shot inference2026.06 | 51.3 | — | — | — | — | — | |
| HuatuoGPT-Vision-7BParameters=7B2025.11 | 50.7 | 53.3 | 70 | 46.7 | 43.3 | 40 | |
| MedGemma-27BParameters=27B2025.11 | 49.3 | 46.7 | 53.3 | 50 | 53.3 | 43.3 | |
| HealthGPT-L14Type=Comprehension only, # Params=14B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 49.2 | — | — | — | — | — | |
| HealthGPT-14BParameters=14B2025.11 | 48 | 50 | 50 | 43.3 | 46.7 | 50 | |
| Qwen3-VLType=Comprehension only, # Params=8B, # Data=-, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 46.7 | — | — | — | — | — | |
| InternVL3.5Type=Comprehension only, # Params=8B, # Data=16.3M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 46 | — | — | — | — | — | |
| InternVL2Type=Comprehension only, # Params=8B, # Data=7.3M, Evaluation Protocol=Zero-shot inference2026.06 | 43.3 | — | — | — | — | — | |
| HuatuoGPT-VisionType=Comprehension only, # Params=7B, # Data=647K, Evaluation Protocol=Zero-shot inference2026.06 | 43.3 | — | — | — | — | — | |
| HealthGPT-M3Type=Comprehension only, # Params=3.8B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 43.3 | — | — | — | — | — | |
| Llama-3.2Type=Comprehension only, # Params=11B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 39.3 | — | — | — | — | — | |
| LLaVA-Med-7BParameters=7B2025.11 | 38.7 | 33.3 | 40 | 26.7 | 40 | 53.3 | |
| Yi-VLType=Comprehension only, # Params=6B, # Data=10K, Evaluation Protocol=Zero-shot inference2026.06 | 38 | — | — | — | — | — | |
| LLaVA-v1.5Type=Comprehension only, # Params=7B, # Data=158K, Evaluation Protocol=Zero-shot inference2026.06 | 32.7 | — | — | — | — | — | |
| OMG-LLaVAType=Unified, # Params=7B, # Data=1.2M, Evaluation Protocol=Zero-shot inference2026.06 | 32.7 | — | — | — | — | — | |
| LLaVA-MedType=Comprehension only, # Params=7B, # Data=60K, Evaluation Protocol=Zero-shot inference2026.06 | 30 | — | — | — | — | — | |
| Med-FlamingoType=Comprehension only, # Params=8.3B, # Data=1.3M, Evaluation Protocol=Zero-shot inference2026.06 | 28.7 | — | — | — | — | — | |
| MedPLIBType=Unified, # Params=14B/7B, # Data=500K, Evaluation Protocol=Zero-shot inference2026.06 | 28.7 | — | — | — | — | — | |
| BLIP-2Type=Comprehension only, # Params=6.7B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 27.3 | — | — | — | — | — | |
| InstructBLIPType=Comprehension only, # Params=7B, # Data=364K, Evaluation Protocol=Zero-shot inference2026.06 | 25.3 | — | — | — | — | — | |
| LLaVA-NeXT-7BParameters=7B2025.11 | 24.7 | 20 | 20 | 26.7 | 33.3 | 23.3 | |
| LLaVA-7BParameters=7B2025.11 | 23.3 | 23.3 | 20 | 26.7 | 23.3 | 23.3 |