Video Referring Segmentation on ReVOS Referring
70.8J&F ScoreVIRST
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VIRSTVenue=-2026.03 | 70.8 | 68.8 | 72.8 | |
| UFVideoSize=7B2025.12 | 67.6 | 65.4 | 69.8 | |
| UniPixelSize=7B2025.12 | 66.4 | 64.2 | 68.5 | |
| Veason-R1MLLM=Qwen2.5VL-7B, Training Protocol=Training-based2025.10 | 63.6 | 60.7 | 66.5 | |
| VRS-HQ-13BCategory=MLLM-based Segmentation Method, Venue=CVPR'252026.03 | 63.3 | 61.1 | 65.5 | |
| VRS-HQ-7BCategory=MLLM-based Segmentation Method, Venue=CVPR'252026.03 | 62.1 | 59.8 | 64.5 | |
| VRS-HQMLLM=ChatUniVi-7B, Training Protocol=Training-based2025.10 | 62.1 | 59.8 | 64.5 | |
| RGA3Size=7B2025.12 | 60.5 | 58.7 | 62.3 | |
| RGA3-7BCategory=MLLM-based Segmentation Method, Venue=ICCV'252026.03 | 60.5 | 58.7 | 62.3 | |
| DecAFMLLM=Qwen2.5VL-7B, Training Protocol=Training-free2025.10 | 58.7 | 54.8 | 62.6 | |
| HyperSeg-3BBackbone=Swin-B, Pre-trained SAM=false2024.11 | 58.5 | 56 | 60.9 | |
| HyperSegSize=3B2025.12 | 58.5 | 56 | 60.9 | |
| HyperSegCategory=MLLM-based Segmentation Method, Venue=CVPR'252026.03 | 58.5 | 56 | 60.9 | |
| GLUSSize=7B2025.12 | 58.3 | 56 | 60.7 | |
| GLUSMLLM=LLaVA-7B, Training Protocol=Training-based2025.10 | 58.3 | 56 | 60.7 | |
| LLaVA-OV-2Model Class=8B2026.05 | 58.2 | — | — | |
| VISASize=13B2025.12 | 57.4 | 55.6 | 59.1 | |
| VISA-13BCategory=MLLM-based Segmentation Method, Venue=ECCV'242026.03 | 57.4 | 55.6 | 59.1 | |
| InstructSegSize=3B2025.12 | 57 | 54.8 | 59.2 | |
| InstructSegCategory=MLLM-based Segmentation Method, Venue=ICCV'252026.03 | 57 | 54.8 | 59.2 | |
| VISA-13BBackbone=ViT-H, Pre-trained SAM=true2024.11 | 54.1 | 52.3 | 55.8 | |
| Loc-HeadMLLM=Qwen2.5VL-7B, Training Protocol=Training-free2025.10 | 53.1 | 49.3 | 56.9 | |
| VISA-7BBackbone=ViT-H, Pre-trained SAM=true2024.11 | 52.9 | 51.1 | 54.7 | |
| Loc-HeadMLLM=Qwen2VL-7B, Training Protocol=Training-free2025.10 | 52.7 | 49.1 | 56.2 | |
| DecAFMLLM=Qwen2VL-7B, Training Protocol=Training-free2025.10 | 52.7 | 48.9 | 56.4 | |
| DecAFMLLM=InternVL3-8B, Training Protocol=Training-free2025.10 | 51.7 | 47.9 | 55.5 | |
| VISA-7BCategory=MLLM-based Segmentation Method, Venue=ECCV'242026.03 | 50.9 | 49.2 | 52.6 | |
| VISAMLLM=ChatUniVi-7B, Training Protocol=Training-based2025.10 | 50.9 | 49.2 | 52.6 | |
| TrackGPT-13BBackbone=ViT-H, Pre-trained SAM=true2024.11 | 49.5 | 48.3 | 50.6 | |
| TrackGPTSize=13B2025.12 | 49.5 | 48.3 | 50.6 | |
| Loc-HeadMLLM=InternVL3-8B, Training Protocol=Training-free2025.10 | 46.7 | 42.9 | 50.6 | |
| LISASize=13B2025.12 | 46.6 | 45.2 | 47.9 | |
| LISA-7BBackbone=ViT-H, Pre-trained SAM=true2024.11 | 45.7 | 44.3 | 47.1 | |
| LISA-7BCategory=MLLM-based Segmentation Method, Venue=CVPR'242026.03 | 45.7 | 44.3 | 47.1 | |
| LISAMLLM=LLaVA-7B, Training Protocol=Training-based2025.10 | 45.7 | 44.3 | 47.1 | |
| DecAFMLLM=LLaVA-OV-7B, Training Protocol=Training-free2025.10 | 43.4 | 39.1 | 47.6 | |
| Loc-HeadMLLM=LLaVA-7B, Training Protocol=Training-free2025.10 | 39.2 | 35 | 43.4 | |
| Qwen3-VLModel Class=8B2026.05 | 37.8 | — | — | |
| LMPMBackbone=Swin-T, Pre-trained SAM=false2024.11 | 34.1 | 29 | 39.1 | |
| LMPMSize=-2025.12 | 34.1 | 29 | 39.1 | |
| LMPMCategory=Segmentation Expert, Venue=ICCV'232026.03 | 34.1 | 29 | 39.1 | |
| Loc-HeadMLLM=LLaVA-OV-7B, Training Protocol=Training-free2025.10 | 32.8 | 29.2 | 36.5 | |
| ReferFormerBackbone=Video-Swin-B, Pre-trained SAM=false2024.11 | 32.7 | 31.2 | 34.3 | |
| ReferFormerSize=-2025.12 | 32.7 | 31.2 | 34.3 | |
| ReferFormerCategory=Segmentation Expert, Venue=CVPR'222026.03 | 32.7 | 31.2 | 34.3 | |
| MTTRSize=-2025.12 | 30 | 29.8 | 30.2 | |
| MTTRCategory=Segmentation Expert, Venue=ECCV'222026.03 | 30 | 29.8 | 30.2 | |
| LLaVA-OV-1.5Model Class=8B2026.05 | 13 | — | — | |
| Keye-VL-1.5Model Class=8B2026.05 | 10.7 | — | — | |
| InternVL-3.5Model Class=8B2026.05 | 10.2 | — | — | |
| PLMModel Class=8B2026.05 | 8.5 | — | — |