Referring Image Segmentation on RefCOCOg (val)
75.3oIoUHIPIE
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| HIPIEBackbone=ViT-H2023.07 | 75.3 | — | — | — | — | — | — | |
| HIPIEBackbone=ViT-H2023.07 | 75.3 | — | — | — | — | — | — | |
| UNINEXTBackbone=ViT-H2023.07 | 74.7 | — | — | — | — | — | — | |
| UNINEXTBackbone=ViT-H2023.07 | 74.7 | — | — | — | — | — | — | |
| GLaMMPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 74.2 | — | — | — | — | — | — | |
| SAM4MLLM-7BPublication=ECCV 2024, Vision Model=SAM-XL, Language Model=Qwen-VL-7B, Multi-dataset training=false2026.02 | 74.2 | — | — | — | — | — | — | |
| GSVA-7BPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 72.7 | — | — | — | — | — | — | |
| SegLLMPublication=ICLR 2025, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 72.6 | — | — | — | — | — | — | |
| Text4Seg-7BPublication=ICLR 2025, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 72.1 | — | — | — | — | — | — | |
| GSVA-7BTraining Strategy=Combined2026.02 | 71.1 | — | — | — | — | — | — | |
| PerceptionGPT-7BTraining Strategy=Combined2026.02 | 70.3 | — | — | — | — | — | — | |
| VIPAPublication=-, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 70.01 | — | — | — | — | 73.6 | — | |
| UNINEXTBackbone=RN502023.07 | 70 | — | — | — | — | — | — | |
| UNINEXTBackbone=RN502023.07 | 70 | — | — | — | — | — | — | |
| HIPIEBackbone=RN502023.07 | 69.8 | — | — | — | — | — | — | |
| HIPIEBackbone=RN502023.07 | 69.8 | — | — | — | — | — | — | |
| PixelLMPublication=CVPR 2024, Vision Model=CLIP-VIT-L, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 69.3 | — | — | — | — | — | — | |
| PixelLM-7BTraining Strategy=Combined2026.02 | 69.3 | — | — | — | — | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 69.2 | 71.15 | — | — | — | — | — | |
| AMLTraining Strategy=Combined2026.02 | 68.84 | — | — | — | — | — | — | |
| LISA-7BPublication=CVPR 2024, Vision Model=SAM-H, Language Model=Vicuna-7B, Multi-dataset training=false2026.02 | 67.9 | — | — | — | — | — | — | |
| MagNetPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 67.79 | — | — | — | — | — | — | |
| MagNetTraining Strategy=Combined2026.02 | 67.79 | — | — | — | — | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 67.76 | 69.36 | — | — | — | — | — | |
| PolyFormer-BPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=true2026.02 | 67.76 | — | — | — | — | 69.36 | — | |
| PolyFormerTraining Strategy=Combined2026.02 | 67.76 | — | — | — | — | — | — | |
| CARIS*Training Strategy=Combined2026.02 | 67.44 | — | — | — | — | — | — | |
| F-LMMPublication=CVPR 2025, Vision Model=SAM-L, Language Model=LLaMA2-7B, Multi-dataset training=false2026.02 | 67.1 | — | — | — | — | — | — | |
| VIPAPublication=-, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 65.93 | — | — | — | — | 69.98 | — | |
| AMLTraining Strategy=Standard2026.02 | 65.67 | — | — | — | — | 69.24 | — | |
| MagNetPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 65.36 | — | — | — | — | — | — | |
| MagNetTraining Strategy=Standard2026.02 | 65.36 | — | — | — | — | 68.53 | — | |
| PGBDTraining Strategy=Standard2026.02 | 65.21 | — | — | — | — | 68.42 | — | |
| CARIS*Training Strategy=Standard2026.02 | 65.15 | — | — | — | — | 68.87 | — | |
| ReLAPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 65 | — | — | — | — | — | — | |
| LQMFormerPublication=CVPR 2024, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 64.73 | — | — | — | — | — | — | |
| RefTRBackbone=RN1012023.07 | 64.7 | — | — | — | — | — | — | |
| RefTRBackbone=RN1012023.07 | 64.7 | — | — | — | — | — | — | |
| CGFormerPublication=CVPR 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 64.68 | — | — | — | — | 67.57 | — | |
| CGFormerTraining Strategy=Standard2026.02 | 64.68 | — | — | — | — | 67.57 | — | |
| NeMoTraining Strategy=Standard2026.02 | 64.4 | — | — | — | — | — | — | |
| ReMamberTraining Strategy=Standard2026.02 | 63.9 | — | — | — | — | — | — | |
| VLTPublication=TPAMI 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 63.49 | — | — | — | — | — | — | |
| DMMIPublication=ICCV 2023, Vision Model=Swin-B, Language Model=BERT-B, Multi-dataset training=false2026.02 | 63.46 | — | — | — | — | — | — | |
| LAVTTraining Strategy=Standard2026.02 | 61.24 | — | — | — | — | 63.34 | — | |
| LTSVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 54.4 | — | — | — | — | — | — | |
| VLTBackbone=Dark562023.07 | 53 | — | — | — | — | — | — | |
| VLTBackbone=Dark562023.07 | 53 | — | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=Step1X-Edit (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 51.23 | — | — | — | — | 54.52 | — | |
| CGANVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 51.01 | — | — | — | — | — | — | |
| MCNVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 49.22 | — | — | — | — | — | — | |
| MAttNetBackbone=RN1012023.07 | 47.6 | — | — | — | — | — | — | |
| MAttNetBackbone=RN1012023.07 | 47.6 | — | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=FLUX.2-klein (9B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=9B2026.05 | 46.9 | — | — | — | — | 50.62 | — | |
| +B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.21 | — | — | — | — | — | — | |
| +B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 46.09 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.87 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.64 | — | — | — | — | — | — | |
| RefAMVision Backbone=DiT, Pretrained Model=SAM, FLUX.1-dev (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 45.53 | — | — | — | — | 47.11 | — | |
| +B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.3 | — | — | — | — | — | — | |
| +B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.26 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.25 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.25 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 45.18 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.99 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.99 | — | — | — | — | — | — | |
| +B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.89 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.76 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.76 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.68 | — | — | — | — | — | — | |
| +B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.63 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.3 | — | — | — | — | — | — | |
| +B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 44.13 | — | — | — | — | — | — | |
| +B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 44.05 | — | — | — | — | — | — | |
| +B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 44.03 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 43.69 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 43.65 | — | — | — | — | — | — | |
| Global-LocalBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 43.29 | — | — | — | — | — | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.92 | — | — | — | — | — | — | |
| Global-LocalBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.85 | — | — | — | — | — | — | |
| +B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.84 | — | — | — | — | — | — | |
| VLM-VGVision Backbone=R101, Pretrained Model=COCO∗, VLM-VG∗, Training protocol=zero-shot methods w/ additional training2026.05 | 42.8 | — | — | — | — | 48 | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.77 | — | — | — | — | — | — | |
| HybridGLVision Backbone=ViT-B, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 42.47 | — | — | — | — | 51.25 | — | |
| +B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.2 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 42.15 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.88 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 41.84 | — | — | — | — | — | — | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.72 | — | — | — | — | — | — | |
| Pseudo-RISVision Backbone=ViT-B, Pretrained Model=SAM, CoCa, CLIP, Training protocol=zero-shot methods w/ additional training2026.05 | 41.63 | — | — | — | — | 45.99 | — | |
| +B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.62 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.52 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.49 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.05 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 40.63 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 40.17 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.4 | — | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 39.37 | — | — | — | — | — | — | |
| +B2GBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 39.24 | — | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.22 | — | — | — | — | — | — |