Referring Segmentation on refCOCO+ (testA) using cIoU
0.842cIoUTraceVision-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| TraceVision-7B2026.02 | 0.842 | |
| HyperSeg-1.5B2026.02 | 0.835 | |
| ReChannelParameters=9B2026.07 | 0.83 | |
| ReChannelParameters=4B2026.07 | 0.816 | |
| FCLMParameters=7B2026.07 | 0.805 | |
| CoPRSParameters=7B2026.07 | 0.803 | |
| UFO (InternVL2.5-8B)Additions=Mask Tokens2026.01 | 0.799 | |
| Youtu-VL (4B)Additions=None2026.01 | 0.796 | |
| UniPixel (Qwne2.5-VL-3B)Additions=SAM Decoder2026.01 | 0.789 | |
| GLaMM (Vicuna-7B)Additions=SAM/Pixel Decoder2026.01 | 0.787 | |
| GLaMMParameters=7B2026.07 | 0.787 | |
| Text4SegParameters=8B2026.07 | 0.779 | |
| Ours (single-object)2026.05 | 0.7672 | |
| Seg-Zero-7BTraining protocol=training-based, Model size=7B2026.05 | 0.762 | |
| Seg-Zero-7B2026.05 | 0.762 | |
| Seg-Agent-7BTraining protocol=training-free, Model size=7B2026.05 | 0.76 | |
| RVG (ViT-B)Additions=MLP Decoder2026.01 | 0.757 | |
| PSALM-3B2026.02 | 0.755 | |
| PSALMParameters=1.3B2026.07 | 0.755 | |
| Qwen2.5-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 0.74 | |
| PerceptionGPT-7BTraining protocol=training-based, Model size=7B2026.05 | 0.739 | |
| PerceptionGPT-7B2026.05 | 0.739 | |
| READ2024.12 | 0.737 | |
| VistaLLM (Vicuna-7B)Additions=None2026.01 | 0.737 | |
| Seg-Zero-3BTraining protocol=training-based, Model size=3B2026.05 | 0.737 | |
| Read-7B2026.05 | 0.737 | |
| READParameters=7B2026.07 | 0.737 | |
| Qwen2.5-VL-7B + SAM2-LTraining protocol=training-free, Model size=7B, SAM model=SAM2-L2026.05 | 0.735 | |
| PixDLM2026.04 | 0.733 | |
| Ours (multi-object)consistent-checkpoint=true2026.05 | 0.733 | |
| Seg-Agent-3BTraining protocol=training-free, Model size=3B2026.05 | 0.732 | |
| OMG-LLaVAParameters=7B2026.07 | 0.731 | |
| SEGLLM2026.05 | 0.73 | |
| PolyFormer2026.07 | 0.729 | |
| PixelLLM-7B2026.02 | 0.721 | |
| PixelLMw/o SAM=true2023.12 | 0.717 | |
| PixelLM-7B2026.02 | 0.717 | |
| PixelLM2026.04 | 0.717 | |
| PixelLM-7BTraining protocol=training-based, Model size=7B2026.05 | 0.717 | |
| PixelLM-7B2026.05 | 0.717 | |
| Qwen2.5-VL-3B + SAM2-LTraining protocol=training-free, Model size=3B, SAM model=SAM2-L2026.05 | 0.715 | |
| SAM3Additions=DETR-like Decoder2026.01 | 0.711 | |
| VisionReasoner-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 0.711 | |
| ReLAw/o SAM=true2023.12 | 0.71 | |
| ReLA2024.12 | 0.71 | |
| ReLA*Training protocol=training-based, Note=traditional approach2026.05 | 0.71 | |
| Seg-R1-7B2026.05 | 0.709 | |
| LISA-7B2024.12 | 0.708 | |
| LISA-7B2026.02 | 0.708 | |
| LISA-7B2026.05 | 0.708 | |
| LISAParameters=7B2026.07 | 0.708 | |
| LaSagnA-7BModel Scale=7B2026.04 | 0.706 | |
| VisionLLM v2 (Swin-T)Additions=Deform-DETR2026.01 | 0.698 | |
| LAVTw/o SAM=true2023.12 | 0.684 | |
| LAVT2024.12 | 0.684 | |
| LAVT*Training protocol=training-based, Note=traditional approach2026.05 | 0.684 | |
| CRISw/o SAM=true2023.12 | 0.681 | |
| CRIS2024.12 | 0.681 | |
| CRIS*Training protocol=training-based, Note=traditional approach2026.05 | 0.681 | |
| CRIS2026.07 | 0.681 | |
| GSVA-7BModel Scale=7B2026.04 | 0.677 | |
| LISAw/o SAM=false2023.12 | 0.674 | |
| LISA-7BModel Scale=7B2026.04 | 0.674 | |
| LISA-7BTraining protocol=training-based, Model size=7B2026.05 | 0.674 | |
| LISAaugw/o SAM=false2023.12 | 0.663 | |
| VLTw/o SAM=true2023.12 | 0.61 | |
| VLT2024.12 | 0.61 | |
| MCNw/o SAM=true2023.12 | 0.55 | |
| MCN2024.12 | 0.55 | |
| Qwen2-VL-7Bre-implemented=true, consistent-checkpoint=true2026.05 | 0.523 | |
| LLaVA-OV-7B2026.05 | 0.471 | |
| RegionVLM-4BZero-shot=true2026.02 | 0.34 |