Surgical Video Question Answering on EndoVis-VQA In-template 18
87.2BLEU-4RL framework over Digital Twin Representations
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RL framework over Digital Twin Representations2026.06 | 87.2 | 91.5 | 88.4 | 59.5 | |
| SurgViVQALLM=GPT-2, Modality=Video, Training Paradigm=Ours2025.11 | 84.94 | 89.65 | 86.08 | 48.27 | |
| SurgViVQA2026.06 | 84.94 | 89.65 | 86.08 | 48.27 | |
| PitVQALLM=GPT-2, Modality=Image, Training Paradigm=Base2025.11 | 81.73 | 86.18 | 83.28 | 40.35 | |
| PitVQA2026.06 | 81.73 | 86.18 | 83.28 | 40.35 | |
| SurgViVQALLM=Qwen, Modality=Video, Training Paradigm=Ours2025.11 | 79.05 | 87.98 | 84.7 | 46.3 | |
| InternVL3 + LoRALLM=-, Modality=Video, Training Paradigm=FT2025.11 | 33.99 | 72.62 | 83.67 | 70.31 | |
| SurgicalGPTLLM=GPT-2, Modality=Image, Training Paradigm=Base2025.11 | 29.55 | 58.6 | 57.99 | 4 | |
| SurgicalGPT2026.06 | 29.55 | 58.6 | 57.99 | 4 | |
| Qwen3-VLParameters=8B2026.06 | 26.3 | 61.4 | 72.5 | 54.6 | |
| Surgical-LVLM2026.06 | 24.6 | 60.8 | 72 | 53.5 | |
| Qwen3-VLParameters=4B2026.06 | 22.5 | 59.3 | 70.2 | 52.8 | |
| Qwen2.5LLM=-, Modality=Image, Training Paradigm=Zero-Shot2025.11 | 19.04 | 57.6 | 68.6 | 51.22 | |
| Qwen2.5-VLParameters=3B2026.06 | 19.04 | 57.6 | 68.6 | 51.22 | |
| MedGemmaLLM=-, Modality=Image, Training Paradigm=Zero-Shot2025.11 | 8.23 | 44.32 | 64.19 | 37.9 | |
| MedGemmaParameters=4B2026.06 | 8.23 | 44.32 | 64.19 | 37.9 | |
| InternVL3LLM=-, Modality=Video, Training Paradigm=Zero-Shot2025.11 | 5.62 | 51.5 | 55.33 | 32.21 | |
| InternVL3Parameters=1B2026.06 | 5.62 | 51.5 | 55.33 | 32.21 | |
| VideoLLaMA3LLM=-, Modality=Video, Training Paradigm=Zero-Shot2025.11 | 3.28 | 52.01 | 54.66 | 40.9 | |
| VideoLLaMA3Parameters=2B2026.06 | 3.28 | 52.01 | 54.66 | 40.9 |