Multi-task Surgical Reasoning and Recognition on SurgCoTBench (test)
64.03Overall ScoreSurgRAW-Qwen3VL-8B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| SurgRAW-Qwen3VL-8BEvaluation Protocol=Zero-Shot, Base Model=Qwen3VL-8B2025.03 | 64.03 | 51.65 | 50.82 | 86.12 | 62.86 | 62.47 | 67.33 | 65.01 | |
| SurgRAW-GPT4oEvaluation Protocol=Zero-Shot, Base Model=GPT-4o2025.03 | 61.73 | 70.48 | 42.87 | 96.21 | 70.39 | 44.52 | 54.13 | 49.81 | |
| Surgical-VQAEvaluation Protocol=Supervised (Training-Required)2025.03 | 47.12 | 50.58 | 49.83 | 73.87 | 58.76 | 38.31 | 34.42 | 30.37 | |
| LLaVa-CoTEvaluation Protocol=Zero-Shot, Base Model=GPT-4o, Technique=Prompt Engineering2025.03 | 38.92 | 34.11 | 23.94 | 62.18 | 40.74 | 35.62 | 33.48 | 34.55 | |
| Qwen3VL-8BEvaluation Protocol=Zero-Shot, Category=General VLM2025.03 | 35.69 | 26.45 | 23.63 | 40.41 | 30.16 | 30.16 | 59.46 | 44.81 | |
| LLaVa-CoTEvaluation Protocol=Zero-Shot, Base Model=Qwen3VL-8B, Technique=Prompt Engineering2025.03 | 34.91 | 18.56 | 44.31 | 52.89 | 38.59 | 8.3 | 52.67 | 31.48 | |
| MDAgentsEvaluation Protocol=Zero-Shot, Base Model=GPT-4o, Framework=Agentic2025.03 | 32.88 | 34.27 | 19.84 | 7.02 | 20.38 | 39.14 | 46.92 | 41.77 | |
| MedAgentsEvaluation Protocol=Zero-Shot, Base Model=GPT-4o, Framework=Agentic2025.03 | 31.96 | 34.1 | 18.35 | 7.88 | 20.11 | 38.02 | 45.4 | 41.71 | |
| GPT-4oEvaluation Protocol=Zero-Shot, Category=General VLM2025.03 | 31.44 | 32.98 | 17.92 | 9.44 | 20.11 | 37.21 | 43.57 | 40.39 | |
| LLava-OV-7BEvaluation Protocol=Zero-Shot, Category=General VLM2025.03 | 27.63 | 19 | 39.63 | 30.23 | 30.74 | 19.16 | 30.88 | 25.12 | |
| MedGemma-4BEvaluation Protocol=Zero-Shot, Category=Domain-adapted Medical VLM2025.03 | 24.17 | 20.28 | 39.74 | 34.01 | 31.91 | 32.85 | 0.33 | 16.59 | |
| LLaVA-Med-7BEvaluation Protocol=Zero-Shot, Category=Domain-adapted Medical VLM2025.03 | 12.46 | 15.39 | 16.69 | 21.94 | 18.01 | 8.44 | 7.42 | 7.93 |