Referring Expression Segmentation on RefCOCOg (test-u)
78.9cIoUHyperSeg
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HyperSegType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 78.9 | — | — | |
| UNINEXTModel Category=Specialist Segmentation Model2024.03 | 76.4 | — | — | |
| UGround-7BParameters=7B2025.10 | 76.1 | — | — | |
| SAM4MLLM-7BType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 75.6 | — | — | |
| POPENFinetuning=true2025.04 | 75.6 | — | — | |
| SAM4MLLM(Qwen)Venue=ECCV'24, Backbone=Qwen, Interactive Model=SAM2025.03 | 75.2 | — | — | |
| SegAgent-Qwen+SClickBackbone=Qwen, Interactive Model=SimpleClick2025.03 | 75.2 | — | — | |
| GLaMMType=MLLM-based Segmentation Network, SAM-based=true2024.11 | 74.9 | — | — | |
| SegAgent-LLaVA+SAMBackbone=LLaVA, Interactive Model=SAM2025.03 | 74.9 | — | — | |
| GLaMM-7BParameters=7B2025.10 | 74.9 | — | — | |
| SegAgent-Qwen+SAMBackbone=Qwen, Interactive Model=SAM2025.03 | 74.62 | — | — | |
| GroundHog-7BType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 74.6 | — | — | |
| POPEN2025.04 | 74.6 | — | — | |
| PSALMType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 74.4 | — | — | |
| SegLLM-7BParameters=7B2025.10 | 73.6 | — | — | |
| GSVA(SAM)Venue=CVPR'24, Interactive Model=SAM2025.03 | 73.3 | — | — | |
| GSVAFinetuning=true2025.04 | 73.3 | — | — | |
| GSVA-7BParameters=7B2025.10 | 73.3 | — | — | |
| OMG-LLaVAType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 72.9 | — | — | |
| OMG-LLaVA2025.10 | 72.9 | — | — | |
| GSVA-7BType=MLLM-based Segmentation Network, SAM-based=true2024.11 | 72 | — | — | |
| GSVA2025.04 | 72 | — | — | |
| LaSagnA-7BType=MLLM-based Segmentation Network, SAM-based=true2024.11 | 71.9 | — | — | |
| PerceptionGPTVenue=CVPR'242025.03 | 71.7 | — | — | |
| READ-7BParameters=7B2025.10 | 71.4 | — | — | |
| SETOKIM2024.06 | 71.3 | — | — | |
| SegAgent-LLaVA+SClickBackbone=LLaVA, Interactive Model=SimpleClick2025.03 | 71.25 | — | — | |
| PolyFormer-LVenue=CVPR'232025.03 | 71.17 | — | — | |
| AnyRefModel Category=Generalist MLLM, Fine-tuned=true2024.03 | 70.7 | — | — | |
| CoRes2025.10 | 70.7 | — | — | |
| LISA-7BModel Category=Generalist MLLM, Fine-tuned=true2024.03 | 70.6 | — | — | |
| LISABackbone=LLaVA-7B2024.07 | 70.6 | — | — | |
| LISA-7BType=MLLM-based Segmentation Network, SAM-based=true2024.11 | 70.6 | — | — | |
| LISA2024.06 | 70.6 | — | — | |
| LISA(SAM)Venue=CVPR'24, Interactive Model=SAM2025.03 | 70.6 | — | — | |
| LISAFinetuning=true2025.04 | 70.6 | — | — | |
| LISA-7BParameters=7B2025.10 | 70.6 | — | — | |
| PixelLM-7BType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 70.5 | — | — | |
| PixelLM2024.06 | 70.5 | — | — | |
| PixelLMVenue=CVPR'242025.03 | 70.5 | — | — | |
| PixelLM2025.04 | 70.5 | — | — | |
| PixelLM-7BParameters=7B2025.10 | 70.5 | — | — | |
| PolyFormerModel Category=Specialist Segmentation Model2024.03 | 70.2 | — | — | |
| AnyRefModel Category=Generalist MLLM, Fine-tuned=false2024.03 | 69.9 | — | — | |
| PolyFormer-BType=Segmentation Specialist, SAM-based=false2024.11 | 69.1 | — | — | |
| Qwen box + SAMInteractive Model=SAM2025.03 | 68.94 | — | — | |
| LISA2025.04 | 68.5 | — | — | |
| LISA-7BModel Category=Generalist MLLM, Fine-tuned=false2024.03 | 68.4 | — | — | |
| NEXT-Chat2024.06 | 67 | — | — | |
| VISABackbone=Chat-UniVi-7B2024.07 | 66.4 | — | — | |
| VISA-7BParameters=7B2025.10 | 66.4 | — | — | |
| GRESModel Category=Specialist Segmentation Model2024.03 | 66 | — | — | |
| ReLABackbone=Swin-B2024.07 | 66 | — | — | |
| ReLA2024.06 | 66 | — | — | |
| ReLA2025.04 | 66 | — | — | |
| ReLA2025.10 | 66 | — | — | |
| LISABackbone=LLaVA-7B, Note=trained via LISA's official GitHub repository2024.07 | 64.8 | — | — | |
| LAVTModel Category=Specialist Segmentation Model2024.03 | 62.1 | — | — | |
| LAVTBackbone=Swin-B2024.07 | 62.1 | — | — | |
| LAVTType=Segmentation Specialist, SAM-based=false2024.11 | 62.1 | — | — | |
| LAVTVenue=CVPR'222025.03 | 62.1 | — | — | |
| LAVT2025.04 | 62.1 | — | — | |
| LAVT2025.10 | 62.1 | — | — | |
| CRISModel Category=Specialist Segmentation Model2024.03 | 60.4 | — | — | |
| CRISBackbone=ResNet1012024.07 | 60.4 | — | — | |
| CRISType=Segmentation Specialist, SAM-based=false2024.11 | 60.4 | — | — | |
| CRISVenue=CVPR'222025.03 | 60.4 | — | — | |
| CRIS2025.04 | 60.4 | — | — | |
| CRIS2025.10 | 60.4 | — | — | |
| Segment Anyword2025.10 | 60.1 | — | — | |
| VLTType=Segmentation Specialist, SAM-based=false2024.11 | 57.7 | — | — | |
| VLT2025.04 | 57.7 | — | — | |
| VLT2025.10 | 57.7 | — | — | |
| VLTBackbone=Darknet532024.07 | 56.7 | — | — | |
| MCNBackbone=Darknet532024.07 | 49.4 | — | — | |
| MCN2025.04 | 49.4 | — | — | |
| MCN2025.10 | 49.4 | — | — | |
| MAttNetVenue=CVPR'182025.03 | 48.61 | — | — | |
| B2GfixBackbone=BLIP2025.09 | — | 25.37 | — | |
| CaRmode=zero-shot2026.05 | — | — | 36.6 | |
| CoHDSize=-2026.01 | — | — | 72.11 | |
| CoHDSize=-2026.01 | — | — | 72.11 | |
| CRIS2023.11 | — | — | 60.4 | |
| CRIS2026.03 | — | — | 60.4 | |
| CRISVenue=CVPR20222024.02 | — | — | 60.4 | |
| GELLA-13B2024.02 | — | — | 71.5 | |
| GELLA-7B2024.02 | — | — | 71.3 | |
| GL-CLIPmode=zero-shot2026.05 | — | — | 42 | |
| GLaMM2023.11 | — | — | 74.9 | |
| Global-Local (+CT)Backbone=BLIP2025.09 | — | 12.52 | — | |
| GRES2023.11 | — | — | 66 | |
| GRESVenue=CVPR20232024.02 | — | — | 66 | |
| GroundhogSize=7B2026.01 | — | — | 74.6 | |
| GSVASize=Vicuna-7B2026.01 | — | — | 73.3 | |
| GSVASize=8B2026.01 | — | — | 73.3 | |
| GSVASize=13B2026.01 | — | — | 77 | |
| HybridGLmode=zero-shot2026.05 | — | — | 51.6 | |
| IteRPrimeEmode=zero-shot2026.05 | — | — | 45.1 | |
| L2Llabeled percentage=0.1%, backbone=Swin Transformer, text_encoder=BERT2026.05 | — | — | 52.8 | |
| LaSagnASize=7B2026.01 | — | — | 71.9 |