Grounded VQA on MedVideoCap In-domain (test)
96.46mIoUGemini 3-Pro-Preview
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini 3-Pro-PreviewModel Class=Proprietary LMM2026.02 | 96.46 | 32.31 | |
| GPT-4oModel Class=Proprietary LMM2026.02 | 88.8 | 29.14 | |
| Gemini-2.5-FlashModel Class=Proprietary LMM2026.02 | 76.74 | 26.96 | |
| MedScope-7B-RLModel Class=Reasoning Agent, Training=RL2026.02 | 61.2 | 39.8 | |
| InternVL3-8BModel Class=Open-Source LMM2026.02 | 57.62 | 26.61 | |
| Qwen2.5-VL-InstructModel Class=Open-Source LMM2026.02 | 56.87 | 25.86 | |
| LongVT-7B-RFTModel Class=Reasoning Agent2026.02 | 55.34 | 31.22 | |
| MedScope-7B-SFTModel Class=Reasoning Agent, Training=SFT2026.02 | 52.7 | 32.2 | |
| VideoLLaMA3-7BModel Class=Open-Source LMM2026.02 | 42.14 | 15.89 | |
| Video-R1-7BModel Class=Reasoning LMM2026.02 | 38.37 | 34.55 | |
| ReWatch-R1-7BModel Class=Reasoning Agent2026.02 | 35.25 | 25.48 | |
| Video-RFTModel Class=Reasoning LMM2026.02 | 31.08 | 28.47 | |
| VideoChat-R1-7BModel Class=Reasoning LMM2026.02 | 23.13 | 33.69 |