Multi-modal Question Answering on MedXpertQA-MM
62.4AccuracyQwen3.5-27B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-27BMode=REASONING, Architecture=Dense, # Total Params=27B, # Activated Params=27B2026.04 | 62.4 | |
| o1Model Category=Proprietary2026.01 | 49.7 | |
| Qwen3-VL-235B-A22BMode=Thinking, Architecture=MoE, # Total Params=236B, # Activated Params=23B2026.04 | 47.6 | |
| EXAONE 4.5 33BMode=REASONING, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 42.1 | |
| Qwen3-VL-32BMode=Thinking, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 41.6 | |
| Gemini2.5-proModel Category=Proprietary2026.01 | 39.5 | |
| PulseMind-72BModel Category=Open-source, Model Scale=~72B2026.01 | 36.7 | |
| GPT-4oDecoding=Greedy2025.08 | 35.95 | |
| MedVLThinker-32B RLDecoding=Greedy2025.08 | 34.6 | |
| GPT-5 miniMode=REASONING: HIGH, Architecture=-, # Total Params=-, # Activated Params=-2026.04 | 34.4 | |
| Lingshu-32BModel Category=Open-source, Model Scale=~32B2026.01 | 30.9 | |
| Gemme 3 27BDecoding=Greedy2025.08 | 30.8 | |
| PulseMind-32BModel Category=Open-source, Model Scale=~32B2026.01 | 29.6 | |
| GPT-4o-miniDecoding=Greedy2025.08 | 28.55 | |
| Qwen2.5-VL-32B-InstructDecoding=Greedy2025.08 | 27.68 | |
| Qwen2.5VL-72BModel Category=Open-source, Model Scale=~72B2026.01 | 27.6 | |
| InternVL3-78BModel Category=Open-source, Model Scale=~72B2026.01 | 27.4 | |
| InternVL3-38BModel Category=Open-source, Model Scale=~32B2026.01 | 25.2 | |
| Qwen2.5VL-32BModel Category=Open-source, Model Scale=~32B2026.01 | 25.2 | |
| MedVLThinker-7B RLDecoding=Greedy2025.08 | 24.43 | |
| SPINEBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 23.84 | |
| MedVLThinker-3B RLDecoding=Greedy2025.08 | 22.9 | |
| TTRLBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 22.61 | |
| Llava Med v1.5 Mistral 7BDecoding=Greedy2025.08 | 22.56 | |
| GPT-4oModel Category=Proprietary2026.01 | 22.3 | |
| LMSIBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 22.01 | |
| HuatuoGPT-Vision-7BDecoding=Greedy2025.08 | 22 | |
| Gemme 3 4BDecoding=Greedy2025.08 | 21.89 | |
| HuatuoGPT-Vision-34BDecoding=Greedy2025.08 | 21.8 | |
| Qwen2.5-VL-3B-InstructDecoding=Greedy2025.08 | 20.69 | |
| Self-ConsistencyBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 19.47 | |
| Qwen2.5-VL-7B-InstructDecoding=Greedy2025.08 | 18.89 | |
| HuatuoGPT-vision-34BModel Category=Open-source, Model Scale=~32B2026.01 | 17.3 | |
| No adaptationBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 17.17 | |
| LLAVA-med-34BModel Category=Open-source, Model Scale=~32B2026.01 | 16.4 | |
| MedGemma 27BDecoding=Greedy2025.08 | 12.13 | |
| MedGemma 4BDecoding=Greedy2025.08 | 8.17 | |
| SEALONGBase Model=Qwen2.5-VL-3B-Instruct2025.11 | 6.51 |