Medical Visual Question Answering on VQA-Med (test)
27.96ROUGE-1Med-Evo
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Med-Evobase_model=Qwen2.5-VL-3B-Instruct2026.03 | 27.96 | — | — | — | — | 57.79 | 14.97 | |
| Base Modelbase_model=Qwen2.5-VL-3B-Instruct2026.03 | 27.45 | — | — | — | — | 56.88 | 12.57 | |
| TTRVbase_model=Qwen2.5-VL-3B-Instruct2026.03 | 25.75 | — | — | — | — | 55.04 | 10.78 | |
| TTRLbase_model=Qwen2.5-VL-3B-Instruct2026.03 | 25.28 | — | — | — | — | 55.96 | 9.58 | |
| EN-INFbase_model=Qwen2.5-VL-3B-Instruct2026.03 | 25.16 | — | — | — | — | 55.96 | 10.18 | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=SFT2025.08 | 24.55 | 22.44 | 95.54 | 90.11 | 58.16 | — | — | |
| ARMed-RModel Category=Fine-tuned VLM2025.08 | 23.17 | 21.51 | 96.56 | 92.36 | 58.4 | — | — | |
| InternVL3-2BModel Category=General VLM2025.08 | 22.82 | 21.44 | 94.93 | 90.43 | 57.41 | — | — | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=GRPO2025.08 | 20.85 | 19.95 | 96.32 | 91.71 | 57.21 | — | — | |
| InternVL3-8BModel Category=General VLM2025.08 | 18.16 | 14.8 | 85.79 | 82.19 | 50.24 | — | — | |
| InternVL3-14BModel Category=General VLM2025.08 | 14.99 | 10.69 | 86.18 | 82.92 | 48.7 | — | — | |
| Qwen2.5-VL-7BModel Category=General VLM2025.08 | 10.4 | 6.64 | 89.53 | 87.36 | 48.48 | — | — | |
| HuatuoGPT-V-7BModel Category=Medical VLM2025.08 | 10.2 | 5.7 | 82.33 | 80.73 | 44.74 | — | — | |
| Qwen2.5-VL-3BModel Category=Fine-tuned VLM, Fine-tuning protocol=Base2025.08 | 6.39 | 4.88 | 69.47 | 67.01 | 36.94 | — | — | |
| LLaVA-v1.6-7BModel Category=General VLM2025.08 | 5.49 | 3.76 | 84.41 | 80.22 | 43.47 | — | — | |
| LLaVA-v1.6-13BModel Category=General VLM2025.08 | 5.47 | 3.76 | 86.08 | 81.41 | 44.18 | — | — | |
| LLaVA-Med-7BModel Category=Medical VLM2025.08 | 3.93 | 2.04 | 86.63 | 84.31 | 44.23 | — | — |