Reasoning Video Object Segmentation on ReVOS Reasoning
61.8J&F ScoreUniPixel
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| UniPixelSize=7B2025.12 | 61.8 | 59.6 | 63.9 | |
| UFVideoSize=7B2025.12 | 61.8 | 59.8 | 63.8 | |
| Veason-R1-7BModel size=7B2025.08 | 59 | 55.8 | 62.2 | |
| Veason-R1MLLM=Qwen2.5VL-7B, Training Protocol=Training-based2025.10 | 59 | 55.8 | 62.2 | |
| Veason-R1Backbone=Qwen2.5VL-7B, Protocol=Fully Trained2026.05 | 59 | 55.8 | 62.2 | |
| RCoT-Seg-3B2026.05 | 58.2 | 54.9 | 61.6 | |
| VRS-HQ-13BModel size=13B2025.08 | 56.8 | 54.1 | 59.4 | |
| Veason-R1-3BModel size=3B2025.08 | 56.8 | 53.6 | 60 | |
| VRS-HQ-13BVenue=CVPR20252026.05 | 56.8 | 54.1 | 59.4 | |
| Veason-R1-3BVenue=CVPR20262026.05 | 56.8 | 53.6 | 60 | |
| VRS-HQ-7BModel size=7B2025.08 | 56.1 | 53.5 | 58.7 | |
| VRS-HQMLLM=ChatUniVi-7B, Training Protocol=Training-based2025.10 | 56.1 | 53.5 | 58.7 | |
| VRS-HQ-7BVenue=CVPR20252026.05 | 56.1 | 53.5 | 58.7 | |
| VRS-HQBackbone=ChatUniVi-7B, Protocol=Fully Trained2026.05 | 56.1 | 53.5 | 58.7 | |
| RGA3Size=7B2025.12 | 55.4 | 53.1 | 57.7 | |
| HyperSegSize=3B2025.12 | 53 | 50.2 | 55.8 | |
| HyperSeg-3BModel size=3B2025.08 | 53 | 50.2 | 55.8 | |
| HyperSeg-3BVenue=CVPR20252026.05 | 53 | 50.2 | 55.8 | |
| SteerSegBackbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 52.4 | 48.9 | 55.9 | |
| InstructSegSize=3B2025.12 | 51.9 | 49.2 | 54.7 | |
| InstructSeg-3BModel size=3B2025.08 | 51.9 | 49.2 | 54.7 | |
| InstructSeg-3BVenue=ICCV20252026.05 | 51.9 | 49.2 | 54.7 | |
| GLUSSize=7B2025.12 | 51.4 | 48.8 | 53.9 | |
| GLUS-7BModel size=7B2025.08 | 51.4 | 48.8 | 53.9 | |
| GLUSMLLM=LLaVA-7B, Training Protocol=Training-based2025.10 | 51.4 | 48.8 | 53.9 | |
| GLUS-7BVenue=CVPR20252026.05 | 51.4 | 48.8 | 53.9 | |
| GLUSBackbone=LLaVA-7B, Protocol=Fully Trained2026.05 | 51.4 | 48.8 | 53.9 | |
| Omni-R1-8BVenue=NeurIPS20252026.05 | 50.7 | — | — | |
| DecAFMLLM=Qwen2.5VL-7B, Training Protocol=Training-free2025.10 | 49.7 | 45.4 | 53.9 | |
| DecAF*Backbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 49.7 | 45.4 | 53.9 | |
| SteerSegBackbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 49.5 | 45.8 | 53.1 | |
| SteerSegBackbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 48.3 | 44.6 | 52 | |
| SteerSegBackbone=LLaVA-OV-7B, Training Protocol=Frozen-LVLM, Training-free=false2026.05 | 47 | 43.4 | 50.6 | |
| TrajSegBackbone=LLaVA-7B2026.03 | 46.1 | 43.7 | 48.5 | |
| VISASize=13B2025.12 | 44.3 | 42 | 46.7 | |
| VISA-13BModel size=13B2025.08 | 44.3 | 42 | 46.7 | |
| VISABackbone=Chat-UniVi-13B2026.03 | 44.3 | 42 | 46.7 | |
| VISA-13BVenue=ECCV20242026.05 | 44.3 | 42 | 46.7 | |
| VISABackbone=LLaVA-13B2026.03 | 44.2 | 41.9 | 46.5 | |
| VISABackbone=LLaVA-7B2026.03 | 43.2 | 40.5 | 45.8 | |
| DecAFMLLM=InternVL3-8B, Training Protocol=Training-free2025.10 | 43.2 | 39.5 | 46.8 | |
| Loc-Head*Backbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 43.2 | 39.5 | 46.8 | |
| DecAF*Backbone=InternVL3-8B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 43.2 | 39.5 | 46.8 | |
| VISA-7BModel size=7B2025.08 | 43 | 40.6 | 45.4 | |
| VISABackbone=Chat-UniVi-7B2026.03 | 43 | 40.6 | 45.4 | |
| VISAMLLM=ChatUniVi-7B, Training Protocol=Training-based2025.10 | 43 | 40.6 | 45.4 | |
| VISA-7BVenue=ECCV20242026.05 | 43 | 40.6 | 45.4 | |
| VISABackbone=ChatUniVi-7B, Protocol=Fully Trained2026.05 | 43 | 40.6 | 45.4 | |
| Loc-HeadMLLM=Qwen2.5VL-7B, Training Protocol=Training-free2025.10 | 40.8 | 37.2 | 44.4 | |
| Loc-Head*Backbone=Qwen2.5VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 40.8 | 37.2 | 44.4 | |
| Loc-HeadMLLM=InternVL3-8B, Training Protocol=Training-free2025.10 | 40.7 | 36.9 | 44.5 | |
| TrackGPTSize=13B2025.12 | 40.5 | 38.1 | 42.9 | |
| TrackGPTBackbone=LLaVA-13B2026.03 | 40.5 | 38.1 | 42.9 | |
| TrackGPTBackbone=LLaVA-7B2026.03 | 39 | 36.8 | 41.2 | |
| DecAFMLLM=Qwen2VL-7B, Training Protocol=Training-free2025.10 | 37.9 | 34.3 | 41.5 | |
| DecAF*Backbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 37.9 | 34.3 | 41.5 | |
| LISASize=13B2025.12 | 36.7 | 34.3 | 39.1 | |
| LISA-13BModel size=13B2025.08 | 36.7 | 34.3 | 39.1 | |
| DecAFMLLM=LLaVA-OV-7B, Training Protocol=Training-free2025.10 | 36.6 | 32.6 | 40.7 | |
| DecAF*Backbone=LLaVA-OV-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 36.6 | 32.6 | 40.7 | |
| LISA-7BModel size=7B2025.08 | 36.1 | 33.8 | 38.4 | |
| LISAMLLM=LLaVA-7B, Training Protocol=Training-based2025.10 | 36.1 | 33.8 | 38.4 | |
| LISABackbone=LLaVA-7B, Protocol=Fully Trained2026.05 | 36.1 | 33.8 | 38.4 | |
| Loc-HeadMLLM=Qwen2VL-7B, Training Protocol=Training-free2025.10 | 35.4 | 32.6 | 38.2 | |
| Loc-Head*Backbone=Qwen2VL-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 35.4 | 32.6 | 38.2 | |
| Loc-HeadMLLM=LLaVA-7B, Training Protocol=Training-free2025.10 | 31.5 | 27.2 | 35.7 | |
| Loc-HeadMLLM=LLaVA-OV-7B, Training Protocol=Training-free2025.10 | 30.6 | 27.4 | 33.7 | |
| Loc-Head*Backbone=LLaVA-7B, Training Protocol=Frozen-LVLM, Training-free=true2026.05 | 28.1 | 23.8 | 32.5 | |
| ReferFormerSize=-2025.12 | 23.4 | 21.3 | 25.6 | |
| ReferFormerBackbone=Video-Swin-B2026.03 | 23.4 | 21.3 | 25.6 | |
| MTTRSize=-2025.12 | 21 | 20.4 | 21.5 | |
| MTTRBackbone=Video-Swin-T2026.03 | 21 | 20.4 | 21.5 | |
| LMPMSize=-2025.12 | 18.8 | 13.3 | 24.3 | |
| LMPM2025.08 | 18.8 | 13.3 | 24.3 | |
| LMPMBackbone=Swin-T2026.03 | 18.8 | 13.3 | 24.3 |