Referring Expression Segmentation on RefCOCO+ (testA)
91.2cIoUEmpirical Upper Bound
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Empirical Upper BoundTraining regime=Oracle reference2026.05 | 91.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlowSegLLM=Qwen3-4B2026.05 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SELF1E-SEG-8Bw/o SMD=true, 1-Token=true, finetuned=false2026.03 | 84.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyperSegType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyperSegw/o SMD=false, 1-Token=false, finetuned=false2026.03 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SELF1E-SEG-2Bw/o SMD=true, 1-Token=true, finetuned=false2026.03 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyperSeg2026.05 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyperSeg2026.05 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3VL-SAMTokSize=4B2026.01 | 83.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RSAgent (ours)Model Version=Qwen2.5-VL, Model Params=7B, RES Train=Y2025.12 | 83 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RICEModel Version=Qwen2.5, Model Params=7B, RES Train=Y2025.12 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RICE2026.05 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HiMTok-8B(ft)w/o SMD=false, 1-Token=false, finetuned=true2026.03 | 82.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SETCON2026.05 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen25VL-SAMTokSize=3B2026.01 | 81.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| u-LLaVAw/o SMD=false, 1-Token=true, finetuned=false2026.03 | 81.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SELF1E-8Bw/o SMD=true, 1-Token=true, finetuned=false2026.03 | 81.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sa2VASize=4B2026.01 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RSAgent-singleModel Version=Qwen2.5, Model Params=7B, RES Train=Y2025.12 | 81.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| X-SAM2026.05 | 81 | — | — | — | — | — | — | — | — | — | — | — | — | |
| X-SAM2026.05 | 81 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UFO-8B(ft)w/o SMD=true, 1-Token=false, finetuned=true2026.03 | 80.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniPixelSize=3B2026.01 | 80.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-Seg4BBackbone=Qwen3-VL-4B2026.05 | 80.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVF-SAMText Prompt Encoder=BEIT-3 (673M), SAM?=true, Training Data=RC, O, A, PP, PIN, HP2024.06 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVF-SAMModel Version=Extra Data, Model Params=-, RES Train=Y2025.12 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLMModel Version=LLaVA1.6, Model Params=8B, RES Train=Y2025.12 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLM-8Bw/o SMD=false, finetuned=false2026.03 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLM2026.05 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVF-SAM2026.05 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UFO-8Bw/o SMD=true, 1-Token=false, finetuned=false2026.03 | 79.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Youtu-VL4BBackbone=VL4B2026.05 | 79.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PaDT ProSize=3B2026.01 | 79.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SELF1E-2Bw/o SMD=true, 1-Token=true, finetuned=false2026.03 | 79.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MLLMSeg2026.05 | 79.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UnipixelQwen2.5-VL-3BBackbone=Qwen2.5-VL-3B2026.05 | 78.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HiMTok-8Bw/o SMD=false, 1-Token=false, finetuned=false2026.03 | 78.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMMType=MLLM-based Segmentation Network, SAM-based=true2024.11 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMMtraining_regime=fine-tuning, model_type=LVLM2025.03 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLAMMText Prompt Encoder=Vicuna (7B), SAM?=true, Training Data=G, RC2024.06 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMMModel Version=Vicuna, Model Params=7B, RES Train=Y2025.12 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMM-7BParameters=7B2025.10 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMMw/o SMD=false, 1-Token=true, finetuned=false2026.03 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MLLMSegInternVL2.5-4BBackbone=InternVL2.5-4B2026.05 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLaMM2026.05 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DETRISModel Version=DETRIS-L, Model Params=-, RES Train=Y2025.12 | 78.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DETRIS2026.05 | 78.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniLSeg-100Text Prompt Encoder=CLIP-B (63M), SAM?=false, Training Data=SA, RC, gRC2024.06 | 78.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVF-SAMText Prompt Encoder=BEIT-3 (673M), SAM?=true, Training Data=RC2024.06 | 78.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniLSegModel Version=UniLSeg-100, Model Params=-, RES Train=Y2025.12 | 78.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniLSeg2026.05 | 78.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Text4Seg (w/ SAM)Architecture Style=Decoder-based2026.01 | 77.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Text4SegInternVL2-8BBackbone=InternVL2-8B2026.05 | 77.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Text4Seg2026.05 | 77.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLM-7BType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 77.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLM-7BLLM=7B2026.05 | 77.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RISEMLLM=Qwen2.5-VL-7B2026.07 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UGround-7BParameters=7B2025.10 | 77.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimpleSegArchitecture Style=Decoder-free, Pre-training status=true, Backbone=Qwen2.5-VL-7B2026.01 | 77.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PinPoint_msmTraining regime=Training-Free2026.05 | 77.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| POPENFinetuning=true2025.04 | 77 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PinPoint_multTraining regime=Training-Free2026.05 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLM(Qwen)Venue=ECCV'24, Backbone=Qwen, Interactive Model=SAM2025.03 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UFOInternVL2-8BBackbone=InternVL2-8B2026.05 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM4MLLMTraining regime=Supervised Fine-tuning2026.05 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegAgent+SAMLLaVA-7BBackbone=LLaVA-7B2026.05 | 76.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegAgent-LLaVA+SAMBackbone=LLaVA, Interactive Model=SAM2025.03 | 76.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| u-LLAVAText Prompt Encoder=Vicuna (7B), SAM?=true, Training Data=A, CS, RC, PL, PV2024.06 | 76.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM-Veteran-7BTraining regime=Reinforcement-Learning Tuning2026.05 | 76.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAM-VeteranMLLM=Qwen2.5-VL-7B2026.07 | 76.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UNINEXTModel Category=Specialist Segmentation Model2024.03 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UNINEXTtraining_regime=fine-tuning2025.03 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HText Prompt Encoder=BERT-B (104M), SAM?=false, Training Data=O, C, RC, V2024.06 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DGSegMLLM=Qwen2.5-VL-7B2026.07 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimpleSegArchitecture Style=Decoder-free, Pre-training status=true, Backbone=Kimi-VL2026.01 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-ZeroModel Version=Qwen2.5, Model Params=7B, RES Train=Y2025.12 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-Zero-7Bstatus=reported scores2026.03 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-Zero-7Bcheckpoint_source=original paper results2026.04 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-Zero2026.05 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-Zero-7BTraining regime=Reinforcement-Learning Tuning2026.05 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-Zero-7BModel Size=7B, setting=zero-shot2025.03 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seg-ZeroMLLM=Qwen2.5-VL-7B2026.07 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegAgent-Qwen+SClickBackbone=Qwen, Interactive Model=SimpleClick2025.03 | 75.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-LVenue=CVPR'232025.03 | 75.71 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-LText Prompt Encoder=BERT-B (104M), SAM?=false, Training Data=RC, gRC2024.06 | 75.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegAgent-Qwen+SAMBackbone=Qwen, Interactive Model=SAM2025.03 | 75.52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALMType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALMtraining_regime=fine-tuning, model_type=LVLM2025.03 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALMText Prompt Encoder=Phi-1.5 (1.3B), SAM?=false, Training Data=C, RC, CI2024.06 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UFOLLaVA-1.5-7BArchitecture Style=Decoder-free2026.01 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALMw/o SMD=false, 1-Token=false, finetuned=false2026.03 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALM2026.05 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PSALM2026.05 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimpleSegArchitecture Style=Decoder-free, Backbone=Kimi-VL2026.01 | 75.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GroundHog-7BType=MLLM-based Segmentation Network, SAM-based=false2024.11 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GroundhogArchitecture Style=Decoder-based2026.01 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GroundHog-7Bw/o SMD=false, 1-Token=false, finetuned=false2026.03 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| WISE-7B-Svariant=shortened2026.04 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Visionreasoner2026.05 | 74.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MIRASTraining Stage=Stage-1, LLaVA version=v1.62025.02 | 74.8 | — | — | — | — | — | — | — | — | — | — | — | — |