Referring Video Object Segmentation on Ref-DAVIS 2017 (val)
80J&FVideoSEG-O3-2B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VideoSEG-O3-2BMethod Category=MLLM-based Methods, Model Scale=2B2026.06 | 80 | 76 | 84 | — | — | — | |
| VIRSTCategory=MLLM-based Segmentation Method, Venue=-2026.03 | 79.5 | 75.9 | 83.1 | — | — | — | |
| VideoSEG-O3-4BMethod Category=MLLM-based Methods, Model Scale=4B2026.06 | 79.4 | 75.5 | 83.2 | — | — | — | |
| UniPixel-7BMethod Category=MLLM-based Methods, Model Scale=7B2026.06 | 76.4 | 72.7 | 80.1 | — | — | — | |
| VRS-HQBackbone=Chat-UniVi-7B2025.01 | 76 | 72.6 | 79.4 | — | — | — | |
| SDAMPublication=Ours2026.03 | 76 | 73.2 | 78.8 | — | — | — | |
| VRS-HQ-7BCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 76 | 72.6 | 79.4 | — | — | — | |
| SAMWISECategory=Segmentation Expert, Venue=CVPR’252026.03 | 74.5 | 70.6 | 67.4 | — | — | — | |
| VRS-HQBackbone=Chat-UniVi-13B2025.01 | 74.4 | 71 | 77.9 | — | — | — | |
| VRS-HQ-13BCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 74.4 | 71 | 77.9 | — | — | — | |
| ViLLa-6BCategory=MLLM-based Segmentation Method, Venue=ICCV’252026.03 | 74.3 | 70.6 | 78 | — | — | — | |
| MPG-SAM2Category=Segmentation Expert, Venue=ICCV’252026.03 | 72.4 | 68.8 | 76 | — | — | — | |
| HyperSegCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 71.2 | — | — | — | — | — | |
| InstructSegCategory=MLLM-based Segmentation Method, Venue=ICCV’252026.03 | 71.1 | 67.3 | 74.9 | — | — | — | |
| InstructSeg-3BMethod Category=MLLM-based Methods, Model Scale=3B2026.06 | 71.1 | 67.3 | 74.9 | — | — | — | |
| GroPromptPublication=-, Referring & Video Training Data=RefC (box), RefYT (box)2024.06 | 70.6 | 67.8 | 73.3 | — | — | — | |
| SAMWISEMethod Category=Specialized Methods2026.06 | 70.6 | 67.4 | 74.5 | — | — | — | |
| VISABackbone=Chat-UniVi-13B2025.01 | 70.4 | 67 | 73.8 | — | — | — | |
| VISAPublication=ECCV 20242026.03 | 70.4 | 67 | 73.8 | — | — | — | |
| VISA-13BCategory=MLLM-based Segmentation Method, Venue=ECCV’242026.03 | 70.4 | 67 | 73.8 | — | — | — | |
| VLP-RVOS (VLMo)Visual Backbone=VLMo-L, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 70.2 | 66.3 | 74.1 | — | — | — | |
| Sa2VA-8BMethod Category=MLLM-based Methods, Model Scale=8B2026.06 | 70 | — | — | — | — | — | |
| RefSAMPublication=arXiv'23, Referring & Video Training Data=RefC, RefYT2024.06 | 69.5 | 65.9 | 73.2 | — | — | — | |
| VideoGLaMMCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 69.5 | 65.6 | 73.3 | — | — | — | |
| VideoGLaMMRetain Tokens=576 (100%)2026.05 | 69.5 | 65.6 | 73.3 | — | — | 100 | |
| VD-ITBackbone=Video Diffusion, Pre-training=RefCOCO/+/g2024.03 | 69.4 | 66.2 | 72.6 | — | — | — | |
| VISABackbone=Chat-UniVi-7B2025.01 | 69.4 | 66.3 | 72.5 | — | — | — | |
| VISA-7BCategory=MLLM-based Segmentation Method, Venue=ECCV’242026.03 | 69.4 | 66.3 | 72.5 | — | — | — | |
| VISA-7BMethod Category=MLLM-based Methods, Model Scale=7B2026.06 | 69.4 | 66.3 | 72.5 | — | — | — | |
| LiteLVLMRetain Tokens=1962026.05 | 69.2 | 66.8 | 71.6 | — | — | 99.5 | |
| ReferDINOCategory=Segmentation Expert, Venue=ICCV’252026.03 | 68.9 | 65.1 | 72.9 | — | — | — | |
| ReferDINOMethod Category=Specialized Methods2026.06 | 68.9 | 65.1 | 72.9 | — | — | — | |
| VideoLISABackbone=LLaVA-Phi-3-V2025.01 | 68.8 | 64.9 | 72.7 | — | — | — | |
| VideoLISAPublication=NeurIPS 20242026.03 | 68.8 | 64.9 | 72.7 | — | — | — | |
| VideoLISACategory=MLLM-based Segmentation Method, Venue=NeurIPS’242026.03 | 68.8 | 64.9 | 72.7 | — | — | — | |
| VideoLISA-3.8BMethod Category=MLLM-based Methods, Model Scale=3.8B2026.06 | 68.8 | 64.9 | 72.7 | — | — | — | |
| VLP-RVOS (CLIP)Visual Backbone=ViT-L/14, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 68.2 | 64.6 | 71.8 | — | — | — | |
| MUTRBackbone=Swin-L2023.05 | 68 | 64.8 | 71.3 | — | — | — | |
| MUTRPublication=arXiv'23, Referring & Video Training Data=RefC, RefYT, AVSB2024.06 | 68 | 64.8 | 71.3 | — | — | — | |
| SSAPublication=CVPR 20252026.03 | 67.3 | 64 | 70.7 | — | — | — | |
| UniRef++-LVisual Encoder=Swin-L2023.12 | 67.2 | 63.4 | 70.9 | — | — | — | |
| UniNEXTPublication=CVPR'23, Referring & Video Training Data=RefC, RefYT, G, La, T, YT, B, V, O2024.06 | 66.7 | 62.3 | 71.1 | — | — | — | |
| ReferDINOSupervision=Text + Mask, Backbone=Video-Swin-T2026.04 | 66.7 | 62.9 | 70.7 | — | — | — | |
| MUTRBackbone=Video-Swin-T2023.05 | 66.5 | 63 | 70 | — | — | — | |
| TrackGPTBackbone=LLaVA-13B2025.01 | 66.5 | 62.7 | 70.4 | — | — | — | |
| MUTRBackbone=Video-Swin-B2023.05 | 66.4 | 62.8 | 70 | — | — | — | |
| DEVAImage Model=ReferFormer, Setting=offline2023.09 | 66.3 | — | — | — | — | — | |
| UniRef-LVisual Encoder=Swin-L2023.12 | 66.3 | 62.9 | 69.7 | — | — | — | |
| DEVAPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT, YT, D, O2024.06 | 66.3 | — | — | — | — | — | |
| UniRefPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT, RefD, YT, O, LV2024.06 | 66.3 | 62.9 | 69.7 | — | — | — | |
| LiteLVLMRetain Tokens=812026.05 | 66.1 | 64.3 | 67.8 | — | — | 95.1 | |
| LISABackbone=LLaVA-13B2025.01 | 66 | 63.2 | 68.8 | — | — | — | |
| LISAPublication=CVPR 20242026.03 | 66 | 63.2 | 68.8 | — | — | — | |
| SOCBackbone=Video-Swin-B, Training Setting=Joint Train, Evaluation Protocol=Zero-shot evaluation from Ref-YouTube-VOS2023.05 | 65.8 | 62.5 | 69.1 | — | — | — | |
| SOCPublication=NeurIPS'23, Referring & Video Training Data=RefC, RefYT2024.06 | 65.8 | 62.5 | 69.1 | — | — | — | |
| HTRVisual Backbone=Swin-L, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 65.6 | 62.3 | 68.8 | — | — | — | |
| VLP-RVOS (VLMo)Visual Backbone=VLMo-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 65.5 | 60.7 | 69.4 | — | — | — | |
| VATEXBackbone=Video-Swin-B2024.04 | 65.4 | 62.3 | 68.5 | — | — | — | |
| MUTRBackbone=ResNet-502023.05 | 65.3 | 62.4 | 68.2 | — | — | — | |
| MUTRBackbone=ResNet-1012023.05 | 65.3 | 61.9 | 68.6 | — | — | — | |
| Grounded-SAMPublication=arXiv'23, Referring & Video Training Data=RefC (box)2024.06 | 65.2 | 62.3 | 68 | — | — | — | |
| DMVSPublication=CVPR 20252026.03 | 65.2 | 62.2 | 68.2 | — | — | — | |
| VLP-RVOS (CLIP)Visual Backbone=ViT-B/16, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 65.1 | 61.4 | 68.8 | — | — | — | |
| DsHmpBackbone=Video-Swin-Base2024.04 | 64.9 | 61.7 | 68.1 | — | — | — | |
| DsHmpVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 64.9 | 61.7 | 68.1 | — | — | — | |
| OnlineReferBackbone=Swin-L, Online/Offline=online2023.07 | 64.8 | 61.6 | 67.7 | — | — | — | |
| OnlineReferPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT2024.06 | 64.8 | 61.6 | 67.7 | — | — | — | |
| OnlineReferVisual Backbone=Swin-L, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 64.8 | 61.6 | 67.7 | — | — | — | |
| OnlineReferBackbone=Swin-L2025.01 | 64.8 | 61.6 | 67.7 | — | — | — | |
| LISABackbone=LLaVA-7B2025.01 | 64.8 | 62.2 | 67.3 | — | — | — | |
| OnlineReferPublication=ICCV 20232026.03 | 64.8 | 61.6 | 67.7 | — | — | — | |
| OnlineReferCategory=Segmentation Expert, Venue=ICCV’232026.03 | 64.8 | 61.6 | 67.7 | — | — | — | |
| LISA-7BCategory=MLLM-based Segmentation Method, Venue=CVPR’242026.03 | 64.8 | 62.2 | 67.3 | — | — | — | |
| LISA-7BMethod Category=MLLM-based Methods, Model Scale=7B2026.06 | 64.8 | 62.2 | 67.3 | — | — | — | |
| TempCDBackbone=Video-Swin-Base2024.04 | 64.6 | 61.6 | 67.6 | — | — | — | |
| TempCDPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT2024.06 | 64.6 | 61.6 | 67.6 | — | — | — | |
| TempCDVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 64.6 | 61.6 | 67.6 | — | — | — | |
| ReferFormerBackbone=Video-Swin-B2023.05 | 64.3 | 60.7 | 68 | — | — | — | |
| LoSh-SBackbone=Video-Swin-B, Number of input frames (w)=52023.06 | 64.3 | 61.8 | 66.8 | — | — | — | |
| VisPrunerRetain Tokens=1962026.05 | 64.3 | 61.2 | 67.4 | — | — | 92.5 | |
| SOCBackbone=Video-Swin-B, Training Setting=With Image Pretrain, Evaluation Protocol=Zero-shot evaluation from Ref-YouTube-VOS2023.05 | 64.2 | 61 | 67.4 | — | — | — | |
| SOCBackbone=Video-Swin-T, Training Setting=Joint Train, Evaluation Protocol=Zero-shot evaluation from Ref-YouTube-VOS2023.05 | 64.2 | 60.9 | 67.5 | — | — | — | |
| SOCBackbone=Video-Swin-Base2024.04 | 64.2 | 61 | 67.4 | — | — | — | |
| SOCVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 64.2 | 61 | 67.4 | — | — | — | |
| DsHmpBackbone=Video-Swin-Tiny2024.04 | 64 | 60.8 | 67.2 | — | — | — | |
| DsHmpSupervision=Text + Mask, Backbone=Video-Swin-T2026.04 | 64 | 60.8 | 67.2 | — | — | — | |
| ReferFormerBackbone=Swin-L2023.05 | 63.9 | 60.8 | 67 | — | — | — | |
| SOCBackbone=Video-Swin-T, Training Setting=With Image Pretrain, Evaluation Protocol=Zero-shot evaluation from Ref-YouTube-VOS2023.05 | 63.5 | 60.2 | 66.7 | — | — | — | |
| UniRef-R50Visual Encoder=ResNet-502023.12 | 63.5 | 60 | 67 | — | — | — | |
| SOCBackbone=Video-Swin-Tiny2024.04 | 63.5 | 60.2 | 66.7 | — | — | — | |
| SOCVisual Backbone=Video-Swin-T, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 63.5 | 60.2 | 66.7 | — | — | — | |
| SOCSupervision=Text + Mask, Backbone=Video-Swin-T2026.04 | 63.5 | 60.2 | 66.7 | — | — | — | |
| SgMgYear=2023, Backbone=Video-Swin-B, Pre-training=RefCOCO/+/g2023.07 | 63.3 | 60.6 | 66 | — | — | — | |
| SgMgBackbone=Video-Swin-B, Number of input frames (w)=52023.06 | 63.3 | 60.6 | 66 | — | — | — | |
| SgMgBackbone=Video-Swin-Base2024.04 | 63.3 | 60.6 | 66 | — | — | — | |
| SgMgPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT2024.06 | 63.3 | 60.6 | 66 | — | — | — | |
| SgMgBackbone=Video-Swin-B, Pre-training=RefCOCO/+/g2024.03 | 63.3 | 60.6 | 66 | — | — | — | |
| SgMgVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 63.3 | 60.6 | 66 | — | — | — | |
| TrackGPTBackbone=LLaVA-7B2025.01 | 63.2 | 59.4 | 67 | — | — | — | |
| VD-ITBackbone=Video Diffusion2024.03 | 63 | 59.9 | 66.1 | — | — | — |