Multi-Modal Visual Question Answering on RAD-ChestCT (val)
44.99AccuracyMed3D-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S2, Reasoning Instruction=w think2026.02 | 44.99 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S2, Reasoning Instruction=w/o think2026.02 | 43.75 | |
| Lingshu-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 35.79 | |
| Lingshu-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 35.19 | |
| Qwen3-VL-4B-ThinkingInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 34.88 | |
| Qwen3-VL-4B-InstructInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 34.48 | |
| Med-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 31.51 | |
| MedGemma-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 30.9 | |
| Qwen3-VL-4B-ThinkingInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 30.85 | |
| Qwen2.5-VL-3B-RLInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 30.71 | |
| Qwen3-VL-4B-InstructInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 30.3 | |
| Med-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 29.99 | |
| Lingshu-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 29.97 | |
| Qwen3-VL-4B-InstructInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 29.42 | |
| Qwen2.5-VL-3BInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 29.34 | |
| Lingshu-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 29.25 | |
| Med-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 29.24 | |
| Qwen3-VL-4B-ThinkingInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 29.07 | |
| MedGemma-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 28.85 | |
| Qwen2.5-VL-3B-RLInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 28.78 | |
| Qwen3-VL-4B-InstructInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 28.74 | |
| MedGemma-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 28.61 | |
| Qwen2.5-VL-3BInput Type=Montage, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 28.48 | |
| Lingshu-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 28.48 | |
| Qwen2.5-VL-3BInput Type=DRR, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 28.36 | |
| Med-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 28.17 | |
| MedGemma-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 27.75 | |
| Qwen2.5-VL-3BInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 27.12 | |
| Qwen3-VL-4B-ThinkingInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 27.06 | |
| Qwen2.5-VL-3BInput Type=Montage, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 26.89 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S1, Reasoning Instruction=w think2026.02 | 26.6 | |
| Qwen2.5-VL-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 26.54 | |
| Lingshu-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 26.46 | |
| Qwen2.5-VL-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 26.35 | |
| MedVLM-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 26.11 | |
| Qwen3-VL-4B-ThinkingInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 25.63 | |
| MedGemma-1.5-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 25.6 | |
| MedVLM-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 25.47 | |
| Qwen2.5-VL-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 25.32 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S1, Reasoning Instruction=w/o think2026.02 | 25.32 | |
| Qwen2.5-VL-3BInput Type=DRR, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 25.17 | |
| Qwen3-VL-4B-InstructInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 25.16 | |
| E3D-GPTInput Type=Volume, LLM Backbone=Vicuna-1.5-7B, Reasoning Instruction=w/o think2026.02 | 25.15 | |
| RadFMInput Type=Volume, LLM Backbone=Llama-2-13B, Reasoning Instruction=w/o think2026.02 | 25 | |
| Qwen2.5-VL-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 24.5 | |
| MedGemma-1.5-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 24.47 | |
| MedGemma-1.5-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 24.4 | |
| MedGemma-1.5-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 24.4 | |
| M3DInput Type=Volume, LLM Backbone=Llama-2-7B, Reasoning Instruction=w/o think2026.02 | 24.28 | |
| Qwen3-VL-4B-ThinkingInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 23.98 | |
| Qwen2.5-VL-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 23.71 | |
| Qwen3-VL-4B-InstructInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 23.46 | |
| Qwen2.5-VL-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 22.28 | |
| MedVLM-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 22.21 | |
| MedVLM-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 22.2 | |
| Med3DVLMInput Type=Volume, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 22.17 | |
| CT-CHATInput Type=Volume, LLM Backbone=Llama-3.1-8B, Reasoning Instruction=w/o think2026.02 | 20.52 |