Medical Visual Question Answering on SLAKE (Open Recall, Closed Accuracy)
90.7Closed AccuracyBioMed-VITAL-13B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BioMed-VITAL-13B2026.03 | 90.7 | 91.69 | |
| CARE-Coord-B2026.03 | 89.19 | 87.34 | |
| FAVP - Vicuna2026.03 | 88.1 | 87.2 | |
| PMC-CLIP2026.03 | 88 | — | |
| InternVL3.5Type=Comprehension only, # Params=8B, # Data=16.3M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 82.1 | 72.2 | |
| Lingshu-7BType=Comprehension only, # Params=7B, # Data=7.1M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 77.2 | 70.8 | |
| Qwen3-VLType=Comprehension only, # Params=8B, # Data=-, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 76.9 | 66.1 | |
| HealthGPT-L14Type=Comprehension only, # Params=14B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 76.4 | 64.5 | |
| MedQwen# Params=7B, Medical VLM=true2026.04 | 75.3 | 59.9 | |
| HealthGPT-M3Type=Comprehension only, # Params=3.8B, # Data=1.5M, Evaluation Protocol=Partially trained on evaluation benchmarks2026.06 | 74.6 | 56.4 | |
| Llama-3.2# Params=11B, Medical VLM=false2026.04 | 72.4 | 37.1 | |
| Llama-3.2Type=Comprehension only, # Params=11B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 72.4 | 52.1 | |
| HealthGPT-L14# Params=14B, Medical VLM=true2026.04 | 71.9 | 56.2 | |
| MEDSIGHTType=Unified, # Params=8B, # Data=72K, Evaluation Protocol=Zero-shot inference2026.06 | 70.9 | 60.2 | |
| HuatuoGPT-VisionType=Comprehension only, # Params=7B, # Data=647K, Evaluation Protocol=Zero-shot inference2026.06 | 69 | 60.1 | |
| InstructBLIP# Params=7B, Medical VLM=false2026.04 | 66.8 | 40.7 | |
| InstructBLIPType=Comprehension only, # Params=7B, # Data=364K, Evaluation Protocol=Zero-shot inference2026.06 | 66.8 | 43.3 | |
| InternVL2# Params=8B, Medical VLM=false2026.04 | 66.6 | 35.2 | |
| InternVL2Type=Comprehension only, # Params=8B, # Data=7.3M, Evaluation Protocol=Zero-shot inference2026.06 | 66.6 | 50.1 | |
| Qwen-2.5-VL# Params=7B, Medical VLM=false2026.04 | 64.7 | 36.7 | |
| HuatuoGPT-Vision# Params=7B, Medical VLM=true2026.04 | 58.5 | 45.6 | |
| LLaVA-MedType=Comprehension only, # Params=7B, # Data=60K, Evaluation Protocol=Zero-shot inference2026.06 | 58.4 | 44.8 | |
| LLaVA-Med# Params=7B, Medical VLM=true2026.04 | 57.7 | 41.3 | |
| MIMO*Type=Unified, # Params=7B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 57 | — | |
| HealthGPT-M3# Params=3.8B, Medical VLM=true2026.04 | 56.4 | 43.6 | |
| OMG-LLaVAType=Unified, # Params=7B, # Data=1.2M, Evaluation Protocol=Zero-shot inference2026.06 | 54.6 | 39.9 | |
| Yi-VL# Params=6B, Medical VLM=false2026.04 | 52.4 | 30.8 | |
| Yi-VLType=Comprehension only, # Params=6B, # Data=10K, Evaluation Protocol=Zero-shot inference2026.06 | 52.4 | 38.4 | |
| MedPLIBType=Unified, # Params=14B/7B, # Data=500K, Evaluation Protocol=Zero-shot inference2026.06 | 48.4 | 37.4 | |
| Med-FlamingoType=Comprehension only, # Params=8.3B, # Data=1.3M, Evaluation Protocol=Zero-shot inference2026.06 | 47 | 25.5 | |
| Med-Flamingo# Params=8.3B, Medical VLM=true2026.04 | 46.4 | 23.8 | |
| BLIP-2# Params=6.7B, Medical VLM=false2026.04 | 41.6 | 32.1 | |
| BLIP-2Type=Comprehension only, # Params=6.7B, # Data=-, Evaluation Protocol=Zero-shot inference2026.06 | 41.6 | 35.3 | |
| LLaVA-v1.5# Params=7B, Medical VLM=false2026.04 | 37.1 | 29.8 | |
| LLaVA-v1.5Type=Comprehension only, # Params=7B, # Data=158K, Evaluation Protocol=Zero-shot inference2026.06 | 37.1 | 37.7 |