Medical Visual Question Answering on SLAKE
90.62AccuracyVisionZip
Evaluation Results
| Method | Links | |
|---|---|---|
| VisionZipBackbone=Lingshu-32B2026.06 | 90.62 | |
| ViToSBackbone=Lingshu-32B2026.06 | 90.62 | |
| VScanBackbone=Lingshu-32B2026.06 | 90.14 | |
| Lingshu-32BModel Category=Open-source, Model Scale=~32B2026.01 | 89.2 | |
| PDropBackbone=Lingshu-32B2026.06 | 88.94 | |
| Hulu-Med-7BReasoning Mode=DirA, Intervention=+ Combined Prior2026.03 | 88.5 | |
| Ours (7B)Parameters=7B2026.03 | 88 | |
| Hulu-Med-7BReasoning Mode=DirA, Intervention=+ Expert Desc.2026.03 | 87.56 | |
| FastVBackbone=Lingshu-32B2026.06 | 87.5 | |
| Hulu-Med-7BReasoning Mode=DirA, Intervention=Baseline2026.03 | 87.37 | |
| Hulu-Med-7BReasoning Mode=DirA, Intervention=+ BBox RoI2026.03 | 87.28 | |
| Lingshu-32BModel Size=>10B, Model Domain=Domain-specific2026.06 | 87.28 | |
| BaselineBackbone=Lingshu-32B2026.06 | 85.82 | |
| Lingshu-7BReasoning Mode=CoT, Intervention=+ Combined Prior2026.03 | 85.77 | |
| PulseMind-72BModel Category=Open-source, Model Scale=~72B2026.01 | 85.6 | |
| Lingshu-7BReasoning Mode=DirA, Intervention=+ Combined Prior2026.03 | 85.57 | |
| Hulu-Med-32BModel Size=>10B, Model Domain=Domain-specific2026.06 | 85.49 | |
| MedVRModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain, Backbone=Qwen2.5-VL-7B2026.04 | 85.3 | |
| InternVL3-8BReasoning Mode=CoT, Intervention=+ Combined Prior2026.03 | 85.29 | |
| InternVL3-8BReasoning Mode=DirA, Intervention=+ Combined Prior2026.03 | 85.11 | |
| Gemini-2.5-proModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 85 | |
| Citrus-VSize=8B2026.03 | 84.9 | |
| ViToSBackbone=Lingshu-7B2026.06 | 84.62 | |
| ViToS2026.06 | 84.6 | |
| Qwen2.5VL 3BAvg. keep rate=1.000, Match / Total=413/4892026.03 | 84.46 | |
| Qwen3VL-8BReasoning Mode=CoT, Intervention=+ Combined Prior2026.03 | 84.45 | |
| Qwen3VL-8BReasoning Mode=DirA, Intervention=+ Combined Prior2026.03 | 84.35 | |
| Photon 3BAvg. keep rate=0.375, Match / Total=412/4892026.03 | 84.25 | |
| Lingshu-7BReasoning Mode=DirA, Intervention=+ BBox RoI2026.03 | 84.07 | |
| OctoMed-7BParameters=7B2026.03 | 84 | |
| IBISAgent2026.01 | 83.5 | |
| MedGemma 4BDecoding=Greedy2025.08 | 83.49 | |
| MedGemma 4BModel Type=Medical LLM2025.10 | 83.49 | |
| Hulu-Med-7BModel Size=<10B, Model Domain=Domain-specific2026.06 | 83.32 | |
| CARE-Flow-BSource Type=Open-source, Parameters=10B2026.03 | 83.21 | |
| CARE-Coord-BSource Type=Open-source2026.03 | 83.11 | |
| LingshuSize=7B2026.03 | 83.1 | |
| Lingshu-7BModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain2026.04 | 83.1 | |
| Lingshu-7BModel Category=Medical Models2026.04 | 83.1 | |
| Phi3.5V-MedEvaluation Protocol=Fine-Tuning, Decoding Strategy=ARCD2025.12 | 83.07 | |
| LINGSHU-7B + MED-SCOUTParameters=7B, Enhancement=Med-Scout2026.01 | 83 | |
| LINGSHU-7BParameters=7B2026.01 | 82.8 | |
| Lingshu-7BReasoning Mode=CoT, Intervention=+ BBox RoI2026.03 | 82.75 | |
| PDropBackbone=Lingshu-7B2026.06 | 82.69 | |
| Lingshu-7BReasoning Mode=DirA, Intervention=+ Expert Desc.2026.03 | 82.37 | |
| Phi3.5V-MedEvaluation Protocol=Fine-Tuning, Decoding Strategy=Greedy2025.12 | 82.28 | |
| Phi3.5V-MedEvaluation Protocol=Fine-Tuning, Decoding Strategy=OPERA2025.12 | 82.28 | |
| Lingshu-32BSource Type=Open-source, Parameters=32B2026.03 | 82.25 | |
| Qwen3VL-8BReasoning Mode=DirA, Intervention=+ Expert Desc.2026.03 | 82 | |
| InternVL3-8BReasoning Mode=DirA, Intervention=+ Expert Desc.2026.03 | 82 | |
| Lingshu-7BReasoning Mode=DirA, Intervention=Baseline2026.03 | 82 | |
| VScanBackbone=Lingshu-7B2026.06 | 81.97 | |
| Phi3.5V-MedEvaluation Protocol=Fine-Tuning, Decoding Strategy=DoLA2025.12 | 81.89 | |
| PulseMind-32BModel Category=Open-source, Model Scale=~32B2026.01 | 81.5 | |
| MEDVISTA-R1Model size=7-13B parameters2026.01 | 81.36 | |
| MEDVISTA-R1Tool=true2026.01 | 81.36 | |
| BaselineBackbone=Lingshu-7B2026.06 | 81.25 | |
| Qwen3VL-8BReasoning Mode=CoT, Intervention=+ Expert Desc.2026.03 | 81.24 | |
| Phi3.5V-MedEvaluation Protocol=Fine-Tuning, Decoding Strategy=VCD2025.12 | 81.1 | |
| VisionZipBackbone=Lingshu-7B2026.06 | 81.01 | |
| ViTAR2026.06 | 80.8 | |
| FastVBackbone=Lingshu-7B2026.06 | 80.53 | |
| InternVL3-8BReasoning Mode=CoT, Intervention=+ Expert Desc.2026.03 | 80.02 | |
| HuatuoGPT-Vision-34BDecoding=Greedy2025.08 | 78.85 | |
| ViToSBackbone=HuatuoGPT-Vision-7B2026.06 | 78.61 | |
| Lingshu-7BReasoning Mode=CoT, Intervention=+ Expert Desc.2026.03 | 78.51 | |
| Lingshu-7BModel Size=<10B, Model Domain=Domain-specific2026.06 | 78.51 | |
| MEDIC-ADSize=7B2026.03 | 78.5 | |
| CARE-Flow-SSource Type=Open-source, Parameters=4B2026.03 | 78.44 | |
| Qwen2.5VL-72BModel Category=Open-source, Model Scale=~72B2026.01 | 78.3 | |
| InternVL3-8BReasoning Mode=DirA, Intervention=+ BBox RoI2026.03 | 78.29 | |
| Lingshu2026.01 | 78 | |
| Qwen3VL-8BReasoning Mode=DirA, Intervention=+ BBox RoI2026.03 | 77.95 | |
| MEDGEMMA-4B-ITParameters=4B2026.01 | 77.9 | |
| Hulu-Med-7BReasoning Mode=CoT, Intervention=+ Combined Prior2026.03 | 77.85 | |
| Lingshu-7BTool=false2026.01 | 77.8 | |
| InternVL3-78BModel Category=Open-source, Model Scale=~72B2026.01 | 77.4 | |
| Chiron-01-8BTool=false2026.01 | 77.4 | |
| MedGemma 27BDecoding=Greedy2025.08 | 77.4 | |
| MedGemma 27BModel Type=Medical LLM2025.10 | 77.4 | |
| Chiron2026.01 | 77.3 | |
| CARE-Coord-SSource Type=Open-source2026.03 | 77.19 | |
| Claude-4.5-haikuModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 77 | |
| GPT-4oDecoding=Greedy2025.08 | 76.44 | |
| MedGemma-4BModel Category=Medical-Specific VLMs, Evaluation Protocol=In-domain2026.04 | 76.4 | |
| GPT-o4-miniModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 76.27 | |
| Lingshu-7BSource Type=Open-source, Parameters=7B2026.03 | 76.15 | |
| Qwen3VL-8BReasoning Mode=CoT, Intervention=+ BBox RoI2026.03 | 76.15 | |
| GEMINI-3-FLASHModel Type=Proprietary2026.01 | 76.1 | |
| Gemini2.5-proModel Category=Proprietary2026.01 | 75.8 | |
| Lingshu-7BReasoning Mode=CoT, Intervention=Baseline2026.03 | 75.77 | |
| InternVL3-8BReasoning Mode=CoT, Intervention=+ BBox RoI2026.03 | 75.68 | |
| QWEN3-VL-4B-INSTRUCT + MED-SCOUTParameters=4B, Enhancement=Med-Scout2026.01 | 75.6 | |
| OpenAI-o3Category=Close-Source SOTA2026.04 | 75.3 | |
| GPT-4o-miniDecoding=Greedy2025.08 | 75.24 | |
| GPT-5.2Model Size=Large, Model Domain=General2026.06 | 75.14 | |
| Qwen3-VL-30B-A3B-InstructModel Size=>10B, Model Domain=General2026.06 | 75.12 | |
| Claude-4.5-sonnetModel size=Proprietary, Prompting protocol=vanilla prompt2026.01 | 75 | |
| HuatuoGPT-Vision-7BDecoding=Greedy2025.08 | 75 | |
| HuatuoGPT-Vision-7BModel Type=Medical LLM2025.10 | 75 |