Referring Video Object Segmentation on Ref-YouTube-VOS (val)
74.2J&F ScoreVIRST
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VIRSTCategory=MLLM-based Segmentation Method, Venue=-2026.03 | 74.2 | — | 72.2 | 76.1 | — | — | |
| MPG-SAM2Category=Segmentation Expert, Venue=ICCV’252026.03 | 73.9 | — | 71.7 | 76.1 | — | — | |
| VideoLoomParameters=8B2026.01 | 71.3 | — | — | — | — | — | |
| VRS-HQBackbone=Chat-UniVi-13B2025.01 | 71 | — | 69 | 73.1 | — | — | |
| VRS-HQParameters=13B2026.01 | 71 | — | — | — | — | — | |
| VRS-HQ-13BCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 71 | — | 69 | 73.1 | — | — | |
| Sa2VA-8BAccess=Specialized open models2026.01 | 70.7 | — | — | — | — | — | |
| Sa2VAParameters=8B2026.01 | 70.7 | — | — | — | — | — | |
| Sa2VA-8BCategory=Specialized open models2026.03 | 70.7 | — | — | — | — | — | |
| SAMWISECategory=Segmentation Expert, Venue=CVPR’252026.03 | 70.6 | — | 69.2 | 67.8 | — | — | |
| MolmoPoint-8BCategory=MolmoPoint2026.03 | 70.5 | — | — | 81.9 | — | 80.5 | |
| VRS-HQBackbone=Chat-UniVi-7B2025.01 | 70.4 | — | 68.3 | 72.5 | — | — | |
| VRS-HQParameters=7B2026.01 | 70.4 | — | — | — | — | — | |
| VRS-HQ-7BCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 70.4 | — | 68.3 | 72.5 | — | — | |
| Molmo2-4BAccess=Open weights, Open data (no distillation), Open code2026.01 | 70.2 | — | — | 80.4 | — | 78.8 | |
| Molmo2-8BAccess=Open weights, Open data (no distillation), Open code2026.01 | 70.2 | — | — | 78.7 | — | 77.3 | |
| Molmo2-4BCategory=Fully open2026.03 | 70.2 | — | — | 80.4 | — | 78.8 | |
| Molmo2-8BCategory=Fully open2026.03 | 70.2 | — | — | 78.7 | — | 77.3 | |
| ReferDINOCategory=Segmentation Expert, Venue=ICCV’252026.03 | 69.3 | — | 67 | 71.5 | — | — | |
| HyperSegCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 68.5 | — | — | — | — | — | |
| MUTRPublication=arXiv'23, Referring & Video Training Data=RefC, RefYT, AVSB2024.06 | 68.4 | — | 66.4 | 70.4 | — | — | |
| Sa2VA-Qwen3-VL-4BAccess=Specialized open models2026.01 | 68.1 | — | — | — | — | — | |
| Sa2VA-Qwen3-VL-4BCategory=Specialized open models2026.03 | 68.1 | — | — | — | — | — | |
| Molmo2-O-7BAccess=Open weights, Open data (no distillation), Open code2026.01 | 67.9 | — | — | 77.7 | — | 76.1 | |
| Molmo2-O-7BCategory=Fully open2026.03 | 67.9 | — | — | 77.7 | — | 76.1 | |
| LMPM++Backbone=Video-Swin-Base2025.12 | 67.8 | — | 65.7 | 69.9 | — | — | |
| VLP-RVOS (VLMo)Visual Backbone=VLMo-L, FLOPS (G)=183, Speed (FPS)=22, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 67.6 | — | 65.3 | 69.8 | — | — | |
| ViLLaParameters=6B2026.01 | 67.5 | — | — | — | — | — | |
| InstructSegCategory=MLLM-based Segmentation Method, Venue=ICCV’252026.03 | 67.5 | — | 65.4 | 69.5 | — | — | |
| ViLLa-6BCategory=MLLM-based Segmentation Method, Venue=ICCV’252026.03 | 67.5 | — | 64.6 | 70.4 | — | — | |
| UniRef-LVisual Encoder=Swin-L2023.12 | 67.4 | — | 65.5 | 69.2 | — | — | |
| UniRefPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT, RefD, YT, O, LV2024.06 | 67.4 | — | 65.5 | 69.2 | — | — | |
| SOCPublication=NeurIPS'23, Referring & Video Training Data=RefC, RefYT2024.06 | 67.3 | — | 65.3 | 69.3 | — | — | |
| SOCLLM usage=no2025.04 | 67.3 | — | 65.3 | 69.3 | — | — | |
| GLUSLLM usage=yes, SFT strategy=additional-SFT2025.04 | 67.3 | — | 65.5 | 69 | — | — | |
| VideoMolmo-7BAccess=Specialized open models2026.01 | 67.3 | — | — | 73.7 | — | — | |
| GLUSParameters=7B2026.01 | 67.3 | — | — | — | — | — | |
| VideoMolmo-7BCategory=Fully open2026.03 | 67.3 | — | — | 73.7 | — | — | |
| LoShBackbone=Video-Swin-Base2025.12 | 67.2 | — | 65.4 | 69 | — | — | |
| DsHmpBackbone=Video-Swin-Base2024.04 | 67.1 | — | 65 | 69.1 | — | — | |
| DsHmpVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 67.1 | — | 65 | 69.1 | — | — | |
| HTRVisual Backbone=Swin-L, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 67.1 | — | 65.3 | 68.9 | — | — | |
| DsHmpLLM usage=no2025.04 | 67.1 | — | 65 | 69.1 | — | — | |
| DsHmpBackbone=Video-Swin-Base2025.12 | 67.1 | — | 65 | 69.1 | — | — | |
| UniRef++-LVisual Encoder=Swin-L2023.12 | 66.9 | — | 64.8 | 69 | — | — | |
| VideoGLaMM2024.11 | 66.8 | — | 65.4 | 68.2 | — | — | |
| VideoGLaMMAccess=Specialized open models2026.01 | 66.8 | — | — | — | — | — | |
| VideoGLaMMCategory=MLLM-based Segmentation Method, Venue=CVPR’252026.03 | 66.8 | — | 65.4 | 68.2 | — | — | |
| VideoGLaMMCategory=Specialized open models2026.03 | 66.8 | — | — | — | — | — | |
| GLUSLLM usage=yes, SFT strategy=standard-SFT2025.04 | 66.6 | — | 65 | 68.3 | — | — | |
| ViLLaLLM usage=yes2025.04 | 66.5 | — | 64.6 | 68.6 | — | — | |
| UniNEXTPublication=CVPR'23, Referring & Video Training Data=RefC, RefYT, G, La, T, YT, B, V, O2024.06 | 66.2 | — | 64 | 68.4 | — | — | |
| DEVAImage Model=ReferFormer, Setting=offline2023.09 | 66 | — | — | — | — | — | |
| SOCBackbone=Video-Swin-Base2024.04 | 66 | — | 64.1 | 67.9 | — | — | |
| DEVAPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT, YT, D, O2024.06 | 66 | — | — | — | — | — | |
| SOCVisual Backbone=Video-Swin-B, FLOPS (G)=98, Speed (FPS)=34, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 66 | — | 64.1 | 67.9 | — | — | |
| VLP-RVOS (CLIP)Visual Backbone=ViT-L/14, FLOPS (G)=219, Speed (FPS)=30, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 66 | — | 63.6 | 68.3 | — | — | |
| SOCBackbone=Video-Swin-Base2025.12 | 66 | — | 64.1 | 67.9 | — | — | |
| TempCDBackbone=Video-Swin-Base2024.04 | 65.8 | — | 63.6 | 68 | — | — | |
| TempCDPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT2024.06 | 65.8 | — | 63.6 | 68 | — | — | |
| TempCDVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 65.8 | — | 63.6 | 68 | — | — | |
| TempCDLLM usage=no2025.04 | 65.8 | — | 63.6 | 68 | — | — | |
| TempCDBackbone=Video-Swin-Base2025.12 | 65.8 | — | 63.6 | 68 | — | — | |
| SgMgYear=2023, Backbone=Video-Swin-B, Pre-training=RefCOCO/+/g2023.07 | 65.7 | — | 63.9 | 67.4 | 40 | — | |
| SgMgBackbone=Video-Swin-Base2024.04 | 65.7 | — | 63.9 | 67.4 | — | — | |
| SgMgPublication=ICCV'23, Referring & Video Training Data=RefC, RefYT2024.06 | 65.7 | — | 63.9 | 67.4 | — | — | |
| SgMgVisual Backbone=Video-Swin-B, FLOPS (G)=121, Speed (FPS)=41, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 65.7 | — | 63.9 | 67.4 | — | — | |
| ReLABackbone=Video-Swin-B2026.01 | 65.7 | — | 63.8 | 67.5 | — | — | |
| SgMgBackbone=Video-Swin-Base2025.12 | 65.7 | — | 63.9 | 67.4 | — | — | |
| GroPromptPublication=-, Referring & Video Training Data=RefC (box), RefYT (box)2024.06 | 65.5 | — | 64.1 | 66.9 | — | — | |
| VATEXBackbone=Video-Swin-B, Pre-training=RefCOCO(+/g) weights2024.04 | 65.4 | — | 63.3 | 67.5 | — | — | |
| SDAMPublication=Ours2026.03 | 65.3 | — | 63.4 | 67.1 | — | — | |
| EPCFormerPublication=arXiv'23, Referring & Video Training Data=RefYT, AVOS2024.06 | 65 | — | 62.9 | 67.2 | — | — | |
| ReferFormerVisual Encoder=Video-Swin2023.12 | 64.9 | — | 62.8 | 67 | — | — | |
| UniLSeg-100variant=1002023.12 | 64.9 | — | 62.8 | 67 | — | — | |
| Molmo + SAM 2Access=Specialized open models2026.01 | 64.6 | — | — | 71.1 | — | — | |
| Molmo + SAM 2Category=Fully open2026.03 | 64.6 | — | — | 71.1 | — | — | |
| DMVSPublication=CVPR 20252026.03 | 64.3 | — | 62.4 | 66.2 | — | — | |
| SSAPublication=CVPR 20252026.03 | 64.3 | — | 62.2 | 66.4 | — | — | |
| ReferFormerVisual Encoder=Swin-L2023.12 | 64.2 | — | 62.3 | 66.2 | — | — | |
| LoShPublication=arXiv'23, Referring & Video Training Data=RefC, RefYT2024.06 | 64.2 | — | 62.5 | 66 | — | — | |
| VATEXBackbone=Swin-L, Pre-training=RefCOCO(+/g) weights2024.04 | 64.2 | — | 61.4 | 67 | — | — | |
| LoShLLM usage=no2025.04 | 64.2 | — | 62.5 | 66 | — | — | |
| UniLSeg-20variant=202023.12 | 64.1 | — | 61.9 | 66.3 | — | — | |
| LMPM++Backbone=Video-Swin-Tiny2025.12 | 64 | — | 62.2 | 65.8 | — | — | |
| VLTSetting=offline2023.09 | 63.8 | — | — | — | — | — | |
| VLT+Visual Encoder=Video-Swin2023.12 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VLTPublication=TPAMI'23, Referring & Video Training Data=RefC, RefYT2024.06 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VLTVisual Backbone=Video-Swin-B, Training Protocol=Pretrained on Ref-COCO/+/g and fine-tuned on Ref-Youtube-VOS2024.05 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VLTBackbone=Video-Swin-B, Pre-training=RefCOCO(+/g) weights2024.04 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VLT2023.12 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VLTBackbone=Video-Swin-B2026.01 | 63.8 | — | 61.9 | 65.6 | — | — | |
| VideoLISABackbone=LLaVA-Phi-3-V2025.01 | 63.7 | — | 61.7 | 65.7 | — | — | |
| VideoLISA2024.11 | 63.7 | — | 65.7 | 61.7 | — | — | |
| VideoLISALLM usage=yes, Model Size=3.8B2025.04 | 63.7 | — | 61.7 | 65.7 | — | — | |
| VideoLISAAccess=Specialized open models2026.01 | 63.7 | — | — | — | — | — | |
| LoShBackbone=Video-Swin-Tiny2025.12 | 63.7 | — | 62 | 65.4 | — | — | |
| VideoLISAPublication=NeurIPS 20242026.03 | 63.7 | — | 61.7 | 65.7 | — | — | |
| VideoLISACategory=MLLM-based Segmentation Method, Venue=NeurIPS’242026.03 | 63.7 | — | 61.7 | 65.7 | — | — | |
| VideoLISACategory=Specialized open models2026.03 | 63.7 | — | — | — | — | — |