Referring Segmentation on refCOCO (val)
86.3cIoUSegCompass-13B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SegCompass-13BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 86.3 | — | |
| HiMTok-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 85.9 | — | |
| SegCompass-7BBackbone=Qwen2.5-VL, Method Category=Interpretable Alignment Method2026.05 | 85.3 | — | |
| X-SAM-3.8BBackbone=Phi-3, Method Category=Latent Query Alignment Method2026.05 | 85.1 | — | |
| ReChannelParameters=9B2026.07 | 84.9 | — | |
| HyperSeg-1.5B2026.02 | 84.8 | — | |
| ReChannelParameters=4B2026.07 | 83.8 | — | |
| PSALM-3B2026.02 | 83.6 | — | |
| PSALMParameters=1.3B2026.07 | 83.6 | — | |
| TraceVision-7B2026.02 | 83.4 | — | |
| FCLMParameters=7B2026.07 | 82.6 | — | |
| LIRA-8BBackbone=InternVL-2, Method Category=Latent Query Alignment Method2026.05 | 81.8 | — | |
| CoPRSParameters=7B2026.07 | 81.6 | — | |
| UFO-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 81 | — | |
| RAS-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 81 | — | |
| UniPixel-7BBackbone=Qwen2.5-VL, Method Category=Latent Query Alignment Method2026.05 | 80.8 | — | |
| Youtu-VL (4B)Additions=None2026.01 | 80.7 | — | |
| UniPixel (Qwne2.5-VL-3B)Additions=SAM Decoder2026.01 | 80.5 | — | |
| SegLLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 80.2 | — | |
| UniRES-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 80.2 | — | |
| Text4Seg-13BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 80.2 | — | |
| UFO (InternVL2.5-8B)Additions=Mask Tokens2026.01 | 80 | — | |
| SegCompass-7BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 80 | — | |
| SAM4MLLM-7BBackbone=LLaVA-1.6, Method Category=Textual Localization Readout Method2026.05 | 79.6 | — | |
| GLaMM (Vicuna-7B)Additions=SAM/Pixel Decoder2026.01 | 79.5 | — | |
| GLaMM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 79.5 | — | |
| GLaMMParameters=7B2026.07 | 79.5 | — | |
| RVG (ViT-B)Additions=MLP Decoder2026.01 | 79.4 | — | |
| Text4Seg-7BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 79.3 | — | |
| SAM-R1Params=7B2026.02 | 79.2 | — | |
| Text4SegParameters=8B2026.07 | 79.2 | — | |
| GroundHog-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 78.5 | — | |
| READ2024.12 | 78.1 | — | |
| READParameters=7B2026.07 | 78.1 | — | |
| OMG-LLaVA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 78 | — | |
| OMG-LLaVAParameters=7B2026.07 | 78 | — | |
| TikArtParams=8B2026.02 | 77.1 | — | |
| PixelLLM-7B2026.02 | 76.9 | — | |
| LaSagnA-7BModel Scale=7B2026.04 | 76.8 | — | |
| LaSagnA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 76.8 | — | |
| VisionLLM v2 (Swin-T)Additions=Deform-DETR2026.01 | 76.6 | — | |
| GSVA-7BModel Scale=7B2026.04 | 76.4 | — | |
| LISA-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 76 | — | |
| SAM3Additions=DETR-like Decoder2026.01 | 75.5 | — | |
| PerceptionGPT-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 75.3 | — | |
| PixDLM2026.04 | 75.2 | — | |
| PerceptionGPT-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 75.1 | — | |
| LISA-7BFine-tuned on ReferSeg=true2023.08 | 74.9 | — | |
| LISA-7B2024.12 | 74.9 | — | |
| LISAParams=13B2026.02 | 74.9 | — | |
| TikArt w/o Zoom ActionParams=8B2026.02 | 74.9 | — | |
| LISA-7B2026.02 | 74.9 | — | |
| LISA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 74.9 | — | |
| LISAParameters=7B2026.07 | 74.9 | — | |
| PolyFormer2026.07 | 74.8 | — | |
| SESAME2024.12 | 74.7 | — | |
| VistaLLM (Vicuna-7B)Additions=None2026.01 | 74.5 | — | |
| SegR1Params=7B2026.02 | 74.3 | — | |
| Seg-R1-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 74.3 | — | |
| LISA-7B2023.08 | 74.1 | — | |
| LISAw/o SAM=false2023.12 | 74.1 | — | |
| LISA-7B2024.08 | 74.1 | — | |
| TikArt w/o ObservationParams=8B2026.02 | 74.1 | — | |
| LISA-7BModel Scale=7B2026.04 | 74.1 | — | |
| LISAaugw/o SAM=false2023.12 | 74 | — | |
| ReLA2023.08 | 73.8 | — | |
| ReLAw/o SAM=true2023.12 | 73.8 | — | |
| GRES2024.08 | 73.8 | — | |
| ReLA2024.12 | 73.8 | — | |
| TikArt w/o ApertureParams=8B2026.02 | 73.8 | — | |
| ReLAMethod Category=Method without LLMs2026.05 | 73.8 | — | |
| FAST2024.08 | 73.3 | — | |
| VPD (UNet)Additions=Denoising Decoder2026.01 | 73.3 | — | |
| PixelLMw/o SAM=true2023.12 | 73 | — | |
| PixelLM-7B2026.02 | 73 | — | |
| PixelLM2026.04 | 73 | — | |
| PixelLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 73 | — | |
| LAVT2023.08 | 72.7 | — | |
| LAVTw/o SAM=true2023.12 | 72.7 | — | |
| LAVT2024.08 | 72.7 | — | |
| LAVT2024.12 | 72.7 | — | |
| LAVTMethod Category=Method without LLMs2026.05 | 72.7 | — | |
| LLaVA w Seg AdapterArchitecture=LLaVA with Segmentation Adapter2024.08 | 70.8 | — | |
| CRIS2023.08 | 70.5 | — | |
| CRISw/o SAM=true2023.12 | 70.5 | — | |
| CRIS2024.12 | 70.5 | — | |
| CRISMethod Category=Method without LLMs2026.05 | 70.5 | — | |
| CRIS2026.07 | 70.5 | — | |
| VLT2023.08 | 67.5 | — | |
| VLTw/o SAM=true2023.12 | 67.5 | — | |
| VLT2024.12 | 67.5 | — | |
| VLTMethod Category=Method without LLMs2026.05 | 67.5 | — | |
| MCN2023.08 | 62.4 | — | |
| MCNw/o SAM=true2023.12 | 62.4 | — | |
| MCN2024.12 | 62.4 | — | |
| RegionVLM-4BZero-shot=true2026.02 | 38.7 | — | |
| AttentionBackbone=Qwen3-VL2026.04 | — | 23 | |
| AttnLRPBackbone=Qwen3-VL2026.04 | — | 27 | |
| Classic Specialist (Non-VLM)Model Category=Classic Specialist, Task-specific fine-tuning=true2026.01 | — | 79.3 | |
| Classic Specialist (VLM)Model Category=Classic Specialist, Task-specific fine-tuning=true2026.01 | — | 80.5 |