Multi-Modal Visual Question Answering on CT-RATE (val)
41.92AccuracyMed3D-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S2, Reasoning Instruction=w think2026.02 | 41.92 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S2, Reasoning Instruction=w/o think2026.02 | 40.3 | |
| Qwen3-VL-4B-InstructInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 33.52 | |
| Qwen3-VL-4B-ThinkingInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 32.77 | |
| Lingshu-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 31.88 | |
| Qwen2.5-VL-3B-RLInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 31.22 | |
| Qwen3-VL-4B-InstructInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 29.93 | |
| Qwen2.5-VL-3BInput Type=Montage, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 29.5 | |
| Med-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 29.32 | |
| Qwen2.5-VL-3BInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 29.3 | |
| Qwen2.5-VL-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 29.23 | |
| Lingshu-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 29.14 | |
| Qwen2.5-VL-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 28.98 | |
| Qwen3-VL-4B-ThinkingInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 28.95 | |
| Qwen2.5-VL-3B-RLInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 28.73 | |
| Qwen3-VL-4B-InstructInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 28.65 | |
| Lingshu-7BInput Type=Montage, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 28.54 | |
| Lingshu-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 28.14 | |
| Qwen3-VL-4B-ThinkingInput Type=Montage, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 28 | |
| Med-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 27.96 | |
| Qwen3-VL-4B-InstructInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 27.86 | |
| Med-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 27.36 | |
| Qwen2.5-VL-3BInput Type=Video, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 27.36 | |
| Qwen2.5-VL-3BInput Type=Montage, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 27.35 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S1, Reasoning Instruction=w/o think2026.02 | 27.32 | |
| Med3D-R1Input Type=Volume, LLM Backbone=Qwen-2.5-3B, Training Stage=S1, Reasoning Instruction=w think2026.02 | 27.12 | |
| MedGemma-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 27.1 | |
| Qwen2.5-VL-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 26.89 | |
| Med-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 26.86 | |
| MedGemma-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 26.76 | |
| Qwen2.5-VL-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 26.63 | |
| MedVLM-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 26.58 | |
| Qwen3-VL-4B-ThinkingInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w/o think2026.02 | 26.54 | |
| Qwen3-VL-4B-ThinkingInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 26.45 | |
| Qwen3-VL-4B-InstructInput Type=Video, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 26.34 | |
| Lingshu-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 26.32 | |
| Lingshu-7BInput Type=Video, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 26.17 | |
| MedVLM-R1Input Type=Montage, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 26.13 | |
| Qwen2.5-VL-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 25.85 | |
| Qwen3-VL-4B-InstructInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 25.85 | |
| RadFMInput Type=Volume, LLM Backbone=Llama-2-13B, Reasoning Instruction=w/o think2026.02 | 25.77 | |
| Qwen2.5-VL-3BInput Type=DRR, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w think2026.02 | 25.7 | |
| MedGemma-1.5-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 25.53 | |
| E3D-GPTInput Type=Volume, LLM Backbone=Vicuna-1.5-7B, Reasoning Instruction=w/o think2026.02 | 25.29 | |
| Qwen2.5-VL-7BInput Type=DRR, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w think2026.02 | 25.01 | |
| MedVLM-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w think2026.02 | 24.19 | |
| MedGemma-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 24.06 | |
| MedGemma-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 24 | |
| Qwen3-VL-4B-ThinkingInput Type=DRR, LLM Backbone=Qwen-3-4B, Reasoning Instruction=w think2026.02 | 23.91 | |
| MedGemma-1.5-4BInput Type=DRR, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 23.35 | |
| CT-CHATInput Type=Volume, LLM Backbone=Llama-3.1-8B, Reasoning Instruction=w/o think2026.02 | 23.14 | |
| Qwen2.5-VL-3BInput Type=DRR, LLM Backbone=Qwen-2.5-3B, Reasoning Instruction=w/o think2026.02 | 23.09 | |
| MedGemma-1.5-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w think2026.02 | 23.07 | |
| M3DInput Type=Volume, LLM Backbone=Llama-2-7B, Reasoning Instruction=w/o think2026.02 | 22.68 | |
| MedVLM-R1Input Type=DRR, LLM Backbone=Qwen-2-2B, Reasoning Instruction=w/o think2026.02 | 22.56 | |
| MedGemma-1.5-4BInput Type=Montage, LLM Backbone=Gemma-3-4B, Reasoning Instruction=w/o think2026.02 | 21.71 | |
| Med3DVLMInput Type=Volume, LLM Backbone=Qwen-2.5-7B, Reasoning Instruction=w/o think2026.02 | 21.59 |