Reasoning Video Object Segmentation on ReasonVOS
75.5J&F ScoreAgentRVOS
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AgentRVOSMLLM=GPT-5, Training Protocol=Training-free Methods2026.03 | 75.5 | 73.1 | 78 | |
| AgentRVOSMLLM=Qwen3-VL-32B-T, Training Protocol=Training-free Methods2026.03 | 70.4 | 67.3 | 73.4 | |
| AgentRVOSMLLM=Qwen3-VL-8B-T, Training Protocol=Training-free Methods2026.03 | 68.6 | 65.5 | 71.8 | |
| SteerSegBackbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 65.9 | 63.1 | 68.7 | |
| CoT-RVSMLLM=GPT-4o, Training Protocol=Training-free Methods2026.03 | 65.5 | 62.4 | 68.7 | |
| DecAF*Backbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 63.9 | 60.5 | 67.2 | |
| SteerSegBackbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 63.6 | 60.8 | 66.4 | |
| SteerSegBackbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 63.3 | 60.6 | 66.1 | |
| Veason-R1Parameters=7B2025.08 | 59.9 | 56 | 63.8 | |
| Veason-R1Backbone=Qwen2.5VL-7B, Training Protocol=Fully Trained2026.05 | 59.9 | 56 | 63.8 | |
| DecAF*Backbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 58.9 | 55.1 | 62.7 | |
| SteerSegBackbone=LLaVA-OV-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 58.6 | 55.7 | 61.5 | |
| RCoT-Seg-3B2026.05 | 58.2 | 54.6 | 61.8 | |
| CoT-RVSMLLM=Qwen3-VL-8B-T, Training Protocol=Training-free Methods, reproduced=true2026.03 | 55.7 | 52.5 | 58.9 | |
| Veason-R1Parameters=3B2025.08 | 55.2 | 51.8 | 58.5 | |
| Veason-R1-3BVenue=CVPR20262026.05 | 55.2 | 51.8 | 58.5 | |
| SDAMPublication=Ours, Reasoning Ability=LLM-based methods with reasoning ability2026.03 | 55.1 | 51.3 | 58.8 | |
| RGA3MLLM=Qwen2.5-VL-7B, Training Protocol=Training-based Methods2026.03 | 53.6 | 51.3 | 56 | |
| DecAF*Backbone=LLaVA-OV-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 52.8 | 49.3 | 56.3 | |
| DecAF*Backbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 52.5 | 49 | 56 | |
| CoT-RVSMLLM=Gemma3-12B, Training Protocol=Training-free Methods2026.03 | 50.7 | 47.5 | 54 | |
| GLUSPublication=CVPR’25, Reasoning Ability=LLM-based methods with reasoning ability2026.03 | 49.9 | 47.5 | 52.4 | |
| GLUSParameters=7B2025.08 | 49.9 | 47.5 | 52.4 | |
| GLUSMLLM=LISA-7B, Training Protocol=Training-based Methods2026.03 | 49.9 | 47.5 | 52.4 | |
| GLUS-7BVenue=CVPR20252026.05 | 49.9 | 47.5 | 52.4 | |
| GLUSBackbone=LLaVA-7B, Training Protocol=Fully Trained2026.05 | 49.9 | 47.5 | 52.4 | |
| VideoLISAPublication=NeurIPS’24, Reasoning Ability=LLM-based methods with reasoning ability2026.03 | 47.5 | 45.1 | 49.9 | |
| VideoLISAParameters=3.8B2025.08 | 47.5 | 45.1 | 49.9 | |
| VideoLISAMLLM=LLaVA-3.8B, Training Protocol=Training-based Methods2026.03 | 47.5 | 45.1 | 49.9 | |
| VideoLISA-3.8BVenue=NeurIPS20242026.05 | 47.5 | 45.1 | 49.9 | |
| VideoLISABackbone=LLaVA-Phi-3-V, Training Protocol=Fully Trained2026.05 | 47.5 | 45.1 | 49.9 | |
| Loc-Head*Backbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 44.3 | 41 | 47.5 | |
| Loc-Head*Backbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 41.1 | 37.9 | 44.3 | |
| OnlineReferPublication=ICCV’23, Reasoning Ability=Traditional methods without reasoning ability2026.03 | 38.7 | 34.6 | 42.9 | |
| OnlineRefer2025.08 | 38.7 | 34.6 | 42.9 | |
| SgMgPublication=ICCV’23, Reasoning Ability=Traditional methods without reasoning ability2026.03 | 36.2 | 33.7 | 38.7 | |
| SgMg2025.08 | 36.2 | 33.7 | 38.7 | |
| SOCPublication=NeurIPS’23, Reasoning Ability=Traditional methods without reasoning ability2026.03 | 35.9 | 33.3 | 38.5 | |
| Loc-Head*Backbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 34 | 31.8 | 36.2 | |
| Loc-Head*Backbone=LLaVA-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 33.6 | 29.3 | 38 | |
| LISAPublication=CVPR’24, Reasoning Ability=LLM-based methods with reasoning ability2026.03 | 31.1 | 29.1 | 33.1 | |
| LISAParameters=7B2025.08 | 31.1 | 29.1 | 33.1 | |
| LISABackbone=LLaVA-7B, Training Protocol=Fully Trained2026.05 | 31.1 | 29.1 | 33.1 |