Referring Segmentation on refCOCO (testA) using cIoU
87.5cIoUSegCompass-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| SegCompass-7BBackbone=Qwen2.5-VL, Method Category=Interpretable Alignment Method2026.05 | 87.5 | |
| SegCompass-13BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 87.3 | |
| X-SAM-3.8BBackbone=Phi-3, Method Category=Latent Query Alignment Method2026.05 | 87.1 | |
| TraceVision-7B2026.02 | 86.8 | |
| ReChannelParameters=9B2026.07 | 86.5 | |
| HiMTok-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 86.3 | |
| HyperSeg-1.5B2026.02 | 85.7 | |
| ReChannelParameters=4B2026.07 | 85.6 | |
| CoPRSParameters=7B2026.07 | 85.3 | |
| PSALM-3B2026.02 | 84.7 | |
| PSALMParameters=1.3B2026.07 | 84.7 | |
| FCLMParameters=7B2026.07 | 83.9 | |
| RAS-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 83.5 | |
| LIRA-8BBackbone=InternVL-2, Method Category=Latent Query Alignment Method2026.05 | 83.4 | |
| GLaMM (Vicuna-7B)Additions=SAM/Pixel Decoder2026.01 | 83.2 | |
| GLaMM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 83.2 | |
| GLaMMParameters=7B2026.07 | 83.2 | |
| UniPixel-7BBackbone=Qwen2.5-VL, Method Category=Latent Query Alignment Method2026.05 | 83 | |
| SegCompass-7BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 82.9 | |
| SAM4MLLM-7BBackbone=LLaVA-1.6, Method Category=Textual Localization Readout Method2026.05 | 82.8 | |
| Text4Seg-13BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 82.7 | |
| UniPixel (Qwne2.5-VL-3B)Additions=SAM Decoder2026.01 | 82.6 | |
| UFO-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 82.6 | |
| Youtu-VL (4B)Additions=None2026.01 | 82 | |
| Text4Seg-7BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 81.9 | |
| UniRES-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 81.8 | |
| Text4SegParameters=8B2026.07 | 81.7 | |
| UFO (InternVL2.5-8B)Additions=Mask Tokens2026.01 | 81.6 | |
| SEGLLM2026.05 | 81.5 | |
| SegLLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 81.5 | |
| RVG (ViT-B)Additions=MLP Decoder2026.01 | 81.2 | |
| Ours (single-object)2026.05 | 80.79 | |
| Seg-Zero-7BTraining protocol=training-based, Model size=7B2026.05 | 80.3 | |
| Seg-Zero-7B2026.05 | 80.3 | |
| OMG-LLaVA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 80.3 | |
| Seg-Zero-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 80.3 | |
| OMG-LLaVAParameters=7B2026.07 | 80.3 | |
| READ2024.12 | 80.2 | |
| PixDLM2026.04 | 80.2 | |
| Read-7B2026.05 | 80.2 | |
| READParameters=7B2026.07 | 80.2 | |
| Seg-Agent-7BTraining protocol=training-free, Model size=7B2026.05 | 79.9 | |
| GroundHog-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 79.9 | |
| VisionLLM v2 (Swin-T)Additions=Deform-DETR2026.01 | 79.3 | |
| Seg-Zero-3BTraining protocol=training-based, Model size=3B2026.05 | 79.3 | |
| SAM-R1-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 79.2 | |
| LISA-7B2024.12 | 79.1 | |
| LISA-7B2026.02 | 79.1 | |
| LISA-7B2026.05 | 79.1 | |
| LISA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 79.1 | |
| PerceptionGPT-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 79.1 | |
| LISAParameters=7B2026.07 | 79.1 | |
| Seg-Agent-3BTraining protocol=training-free, Model size=3B2026.05 | 79 | |
| VisionReasoner-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 78.9 | |
| LISA-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 78.8 | |
| LaSagnA-7BModel Scale=7B2026.04 | 78.7 | |
| Seg-R1-7B2026.05 | 78.7 | |
| LaSagnA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 78.7 | |
| Seg-R1-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 78.7 | |
| PerceptionGPT-7BTraining protocol=training-based, Model size=7B2026.05 | 78.6 | |
| PerceptionGPT-7B2026.05 | 78.6 | |
| PerceptionGPT-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 78.6 | |
| PixelLLM-7B2026.02 | 78.5 | |
| Ours (multi-object)consistent-checkpoint=true2026.05 | 78.04 | |
| Qwen2.5-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 77.9 | |
| Qwen2.5-VL-7B + SAM2-LTraining protocol=training-free, Model size=7B, SAM model=SAM2-L2026.05 | 77.8 | |
| SAM3Additions=DETR-like Decoder2026.01 | 77.6 | |
| GSVA-7BModel Scale=7B2026.04 | 77.4 | |
| VisionReasoner-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 77.4 | |
| PolyFormer2026.07 | 76.6 | |
| ReLA2024.12 | 76.5 | |
| PixelLM-7B2026.02 | 76.5 | |
| LISA-7BModel Scale=7B2026.04 | 76.5 | |
| PixelLM2026.04 | 76.5 | |
| ReLA*Training protocol=training-based, Note=traditional approach2026.05 | 76.5 | |
| LISA-7BTraining protocol=training-based, Model size=7B2026.05 | 76.5 | |
| PixelLM-7BTraining protocol=training-based, Model size=7B2026.05 | 76.5 | |
| PixelLM-7B2026.05 | 76.5 | |
| ReLAMethod Category=Method without LLMs2026.05 | 76.5 | |
| PixelLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 76.5 | |
| VistaLLM (Vicuna-7B)Additions=None2026.01 | 76 | |
| Qwen2.5-VL-3B + SAM2-LTraining protocol=training-free, Model size=3B, SAM model=SAM2-L2026.05 | 75.9 | |
| LAVT2024.12 | 75.8 | |
| LAVT*Training protocol=training-based, Note=traditional approach2026.05 | 75.8 | |
| LAVTMethod Category=Method without LLMs2026.05 | 75.8 | |
| CRIS2024.12 | 73.2 | |
| CRIS*Training protocol=training-based, Note=traditional approach2026.05 | 73.2 | |
| CRISMethod Category=Method without LLMs2026.05 | 73.2 | |
| CRIS2026.07 | 73.2 | |
| VLT2024.12 | 70.5 | |
| VLTMethod Category=Method without LLMs2026.05 | 70.5 | |
| MCN2024.12 | 64.2 | |
| Qwen2-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 58.9 | |
| LLaVA-OV-7B2026.05 | 58.1 | |
| RegionVLM-4BZero-shot=true2026.02 | 39.4 |