Referring Image Segmentation on RefCOCOg (test)
74.9oIoUGLaMM
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GLaMMPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 74.9 | — | — | — | — | — | — | |
| SAM4MLLM-7BPublication=ECCV 2024, Vision Model=SAM-XL, Language Model=Qwen-VL-7B, Multi-dataset training=false2026.02 | 74.3 | — | — | — | — | — | — | |
| Text4Seg-7BPublication=ICLR 2025, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 73.9 | — | — | — | — | — | — | |
| SegLLMPublication=ICLR 2025, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 73.6 | — | — | — | — | — | — | |
| GSVA-7BPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 73.3 | — | — | — | — | — | — | |
| GSVA-7BTraining Strategy=Combined2026.02 | 72 | — | — | — | — | — | — | |
| VIPAPublication=-, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 71.77 | — | — | — | — | 74.97 | — | |
| PerceptionGPT-7BTraining Strategy=Combined2026.02 | 71.7 | — | — | — | — | — | — | |
| LISA-7BPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 70.6 | — | — | — | — | — | — | |
| PixelLMPublication=CVPR 2024, Vision Model=CLIP-VIT-L, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 70.5 | — | — | — | — | — | — | |
| PixelLM-7BTraining Strategy=Combined2026.02 | 70.5 | — | — | — | — | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 70.19 | 71.17 | — | — | — | — | — | |
| AMLTraining Strategy=Combined2026.02 | 70.01 | — | — | — | — | — | — | |
| MagNetPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 69.29 | — | — | — | — | — | — | |
| MagNetTraining Strategy=Combined2026.02 | 69.29 | — | — | — | — | — | — | |
| CARIS*Training Strategy=Combined2026.02 | 69.17 | — | — | — | — | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 69.05 | 69.88 | — | — | — | — | — | |
| PolyFormer-BPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 69.05 | — | — | — | — | 69.88 | — | |
| PolyFormerTraining Strategy=Combined2026.02 | 69.05 | — | — | — | — | — | — | |
| VIPAPublication=-, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 67.37 | — | — | — | — | 70.94 | — | |
| AMLTraining Strategy=Standard2026.02 | 66.78 | — | — | — | — | 69.73 | — | |
| PGBDTraining Strategy=Standard2026.02 | 66.31 | — | — | — | — | 67.99 | — | |
| VLTPublication=TPAMI 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 66.22 | — | — | — | — | — | — | |
| LQMFormerPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 66.04 | — | — | — | — | — | — | |
| MagNetPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 66.03 | — | — | — | — | — | — | |
| MagNetTraining Strategy=Standard2026.02 | 66.03 | — | — | — | — | 69.15 | — | |
| ReLAPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 65.97 | — | — | — | — | — | — | |
| CGFormerPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 65.09 | — | — | — | — | 67.83 | — | |
| CGFormerTraining Strategy=Standard2026.02 | 65.09 | — | — | — | — | 67.83 | — | |
| CARIS*Training Strategy=Standard2026.02 | 65 | — | — | — | — | 68.51 | — | |
| NeMoTraining Strategy=Standard2026.02 | 64.8 | — | — | — | — | — | — | |
| DMMIPublication=ICCV 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 64.19 | — | — | — | — | — | — | |
| ReMamberTraining Strategy=Standard2026.02 | 64 | — | — | — | — | — | — | |
| LAVTTraining Strategy=Standard2026.02 | 62.09 | — | — | — | — | 63.62 | — | |
| LTSVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 54.25 | — | — | — | — | — | — | |
| CGANVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 51.69 | — | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=Step1X-Edit (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 50.86 | — | — | — | — | 54.25 | — | |
| MCNVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 49.4 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.34 | — | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=FLUX.2-klein (9B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=9B2026.05 | 46.26 | — | — | — | — | 50.47 | — | |
| +B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.24 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.21 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.09 | — | — | — | — | — | — | |
| +B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.03 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.9 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.6 | — | — | — | — | — | — | |
| RefAMVision Backbone=DiT, Pretrained Model=SAM, FLUX.1-dev (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 44.45 | — | — | — | — | 48.35 | — | |
| +B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.4 | — | — | — | — | — | — | |
| VLM-VGVision Backbone=R101, Pretrained Model=COCO∗, VLM-VG∗, Training protocol=zero-shot methods w/ additional training2026.05 | 44.1 | — | — | — | — | 48.5 | — | |
| +B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.06 | — | — | — | — | — | — | |
| +B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 44.02 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 43.8 | — | — | — | — | — | — | |
| Pseudo-RISVision Backbone=ViT-B, Pretrained Model=SAM, CoCa, CLIP, Training protocol=zero-shot methods w/ additional training2026.05 | 43.52 | — | — | — | — | 46.67 | — | |
| HybridGLVision Backbone=ViT-B, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 42.97 | — | — | — | — | 51.59 | — | |
| Global-LocalBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.96 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 42.79 | — | — | — | — | — | — | |
| +B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.58 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.31 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.19 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.93 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.46 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 40.07 | — | — | — | — | — | — | |
| Global-LocalBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.24 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 38.94 | — | — | — | — | — | — | |
| +B2GBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 38.9 | — | — | — | — | — | — | |
| Ref-DiffVision Backbone=ViT-B, Pretrained Model=SAM, SD, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 37.5 | — | — | — | — | 44.51 | — | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 36.51 | — | — | — | — | — | — | |
| TASVision Backbone=ViT-B, Pretrained Model=SAM, BLIP2, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 36.16 | — | — | — | — | 46.8 | — | |
| SAM 3Vision Backbone=PE, Pretrained Model=SAM 3, Training protocol=zero-shot methods w/o additional training2026.05 | 35.81 | — | — | — | — | 32.08 | — | |
| Global-LocalVision Backbone=R50, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 30.48 | — | — | — | — | 40.94 | — | |
| Global-LocalVision Backbone=R50, Pretrained Model=FreeSOLO, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 29.83 | — | — | — | — | 33.12 | — | |
| Global-LocalVision Backbone=ViT-B, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 28.21 | — | — | — | — | 40.74 | — | |
| Grad-CAMVision Backbone=R50, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 23.91 | — | — | — | — | 32.5 | — | |
| MaskCLIPVision Backbone=R50, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 23.41 | — | — | — | — | 30.15 | — | |
| B2GBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 47.2 | — | |
| B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 50.74 | — | |
| B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 53.62 | — | |
| B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.23 | — | |
| B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 51.9 | — | |
| B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.98 | — | |
| B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 53.76 | — | |
| B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.45 | — | |
| B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 51.43 | — | |
| B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.94 | — | |
| B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 53.84 | — | |
| B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.62 | — | |
| B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 51.8 | — | |
| B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | — | — | — | — | — | 55.88 | — | |
| CaRVision Backbone=ViT-B and ViT-L, Pretrained Model=CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | — | — | — | — | — | 36.57 | — | |
| CGANBackbone=DarkNet53, DCRF=false2021.03 | — | 51.69 | — | — | — | — | — | |
| CGFormerGeneralization setting=True2026.02 | — | — | — | 65.67 | 42.31 | — | — | |
| CGFormerImage Encoder=Swin-B, Text Encoder=BERT2026.03 | — | — | 65.09 | — | — | — | 67.83 | |
| CRISVisual Backbone=RN101, Text Encoder=GPT-22023.02 | — | 60.36 | — | — | — | — | — | |
| CRISGeneralization setting=True2026.02 | — | — | — | 59.68 | 38.88 | — | — | |
| CRISTraining Strategy=Standard2026.02 | — | — | — | — | — | 60.36 | — | |
| CRISImage Encoder=ResNet-101, Text Encoder=GPT-22026.03 | — | — | — | — | — | — | 60.36 | |
| DETRISImage Encoder=DINOv2-B, Text Encoder=CLIP2026.03 | — | — | — | — | — | — | 68.1 | |
| DETRIS-BPublication=AAAI 2025, Vision Model=DINOv2-B, Language Model=CLIP, Multi-dataset training=false2026.02 | — | — | — | — | — | 68.1 | — | |
| EEVGPublication=ECCV 2024, Vision Model=ViT-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | — | — | — | — | — | 70.01 | — |