Referring Segmentation on refCOCOg (val)
84CIoUSegCompass-13B
Evaluation Results
| Method | Links | |
|---|---|---|
| SegCompass-13BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 84 | |
| X-SAM-3.8BBackbone=Phi-3, Method Category=Latent Query Alignment Method2026.05 | 83.8 | |
| SegCompass-7BBackbone=Qwen2.5-VL, Method Category=Interpretable Alignment Method2026.05 | 82.8 | |
| HiMTok-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 80.1 | |
| HyperSeg-1.5B2026.02 | 79.4 | |
| SegCompass-7BBackbone=LLaVA-1.5, Method Category=Interpretable Alignment Method2026.05 | 79.4 | |
| LIRA-8BBackbone=InternVL-2, Method Category=Latent Query Alignment Method2026.05 | 78.4 | |
| TraceVision-7B2026.02 | 77.6 | |
| UFO-8BBackbone=InternVL-2.5, Method Category=Latent Query Alignment Method2026.05 | 76.7 | |
| Youtu-VL (4B)Additions=None2026.01 | 76.5 | |
| UniPixel-7BBackbone=Qwen2.5-VL, Method Category=Latent Query Alignment Method2026.05 | 76.4 | |
| UniPixel (Qwne2.5-VL-3B)Additions=SAM Decoder2026.01 | 76.3 | |
| RAS-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 76 | |
| UFO (InternVL2.5-8B)Additions=Mask Tokens2026.01 | 75.5 | |
| SAM4MLLM-7BBackbone=LLaVA-1.6, Method Category=Textual Localization Readout Method2026.05 | 74.5 | |
| Ours (single-object)2026.05 | 74.26 | |
| GLaMM (Vicuna-7B)Additions=SAM/Pixel Decoder2026.01 | 74.2 | |
| GLaMM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 74.2 | |
| GroundHog-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 74.1 | |
| Text4Seg-13BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 74 | |
| PSALM-3B2026.02 | 73.8 | |
| UniRES-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 73.8 | |
| SEGLLM2026.05 | 73.6 | |
| SAM3Additions=DETR-like Decoder2026.01 | 73.4 | |
| PixDLM2026.04 | 73.3 | |
| OMG-LLaVA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 72.9 | |
| Seg-Zero-7B2026.05 | 72.6 | |
| SegLLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 72.6 | |
| Text4Seg-7BBackbone=LLaVA-1.5, Method Category=Textual Localization Readout Method2026.05 | 72.1 | |
| PerceptionGPT-7B2026.05 | 71.7 | |
| Read-7B2026.05 | 71.4 | |
| Seg-R1-7B2026.05 | 71.4 | |
| RVG (ViT-B)Additions=MLP Decoder2026.01 | 71.3 | |
| GSVA-7BModel Scale=7B2026.04 | 71.1 | |
| Ours (multi-object)consistent-checkpoint=true2026.05 | 71.08 | |
| Seg-R1-7BBackbone=Qwen2.5-VL, Method Category=Textual Localization Readout Method2026.05 | 71 | |
| Qwen2.5-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 70.9 | |
| VisionLLM v2 (Swin-T)Additions=Deform-DETR2026.01 | 70.7 | |
| PixelLLM-7B2026.02 | 70.7 | |
| PerceptionGPT-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 70.7 | |
| LaSagnA-7BModel Scale=7B2026.04 | 70.6 | |
| LISA-7B2026.05 | 70.6 | |
| LaSagnA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 70.6 | |
| PixelLM-7B2026.05 | 70.5 | |
| PerceptionGPT-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 70.3 | |
| LISA-13BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 69.5 | |
| PixelLM-7B2026.02 | 69.3 | |
| PixelLM2026.04 | 69.3 | |
| PixelLM-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 69.3 | |
| VistaLLM (Vicuna-7B)Additions=None2026.01 | 69 | |
| VisionReasoner-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 68.3 | |
| LISA-7B2026.02 | 67.9 | |
| LISA-7BBackbone=LLaVA-1.5, Method Category=Latent Query Alignment Method2026.05 | 67.9 | |
| FAST2024.08 | 67 | |
| LISA-7B2024.08 | 66.4 | |
| SEEM2024.08 | 65.7 | |
| SEEMMethod Category=Method without LLMs2026.05 | 65.7 | |
| SEEM (DaViT-d5)Additions=SEEM-Decoder2026.01 | 65.6 | |
| GRES2024.08 | 65 | |
| ReLAMethod Category=Method without LLMs2026.05 | 65 | |
| X-Decoder2024.08 | 64.6 | |
| X-DecoderMethod Category=Method without LLMs2026.05 | 64.6 | |
| LISA-7BModel Scale=7B2026.04 | 64.5 | |
| LLaVA w Seg AdapterArchitecture=LLaVA with Segmentation Adapter2024.08 | 64 | |
| VPD (UNet)Additions=Denoising Decoder2026.01 | 62 | |
| LAVT2024.08 | 61.2 | |
| LAVTMethod Category=Method without LLMs2026.05 | 61.2 | |
| CRISMethod Category=Method without LLMs2026.05 | 59.9 | |
| LLaVA-OV-7B2026.05 | 55.6 | |
| VLTMethod Category=Method without LLMs2026.05 | 55 | |
| Qwen2-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 52.1 | |
| RegionVLM-4BZero-shot=true2026.02 | 33.9 |