Referring Image Segmentation on RefCOCO+ (testA)
2,982mIoUCAFT
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CAFTZero-shot=true2026.02 | 2,982 | — | — | — | — | — | |
| SaGZero-shot=true2026.02 | 1,986 | — | — | — | — | — | |
| FLAIRZero-shot=true2026.02 | 1,825 | — | — | — | — | — | |
| MCLIPZero-shot=true2026.02 | 1,201 | — | — | — | — | — | |
| GViTZero-shot=true2026.02 | 679 | — | — | — | — | — | |
| DeRIS-L (published)Visual Enc.=Swin-B, Text Enc.=BEiT3-L, Evaluation pipeline=Reference (published protocol)2026.06 | 83.74 | — | — | — | — | — | |
| Venice-H1Visual Enc.=Swin-B, Text Enc.=BEiT3-L, Evaluation pipeline=Our evaluation pipeline2026.06 | 83.46 | — | — | — | — | 0.08 | |
| DeRIS-L (reproduced)Visual Enc.=Swin-B, Text Enc.=BEiT3-L, Evaluation pipeline=Our evaluation pipeline2026.06 | 83.38 | — | — | — | — | — | |
| OneRef-LVisual Enc.=BEiT3-L, Text Enc.=BEiT3-L2026.06 | 80.16 | — | — | — | — | — | |
| C3VGVisual Enc.=BEiT3-B, Text Enc.=BEiT3-B2026.06 | 79.61 | — | — | — | — | — | |
| PPCRMLLMs=LLaVA-7B2026.03 | 79.56 | — | — | 78.52 | — | — | |
| DeRIS-BVisual Enc.=Swin-S, Text Enc.=BEiT3-B2026.06 | 79.16 | — | — | — | — | — | |
| DETRIS-L*Tuning Protocol=Parameter Efficient Tuning, Training Data Strategy=Mixed RefCOCO dataset, Model Scale=Large2025.01 | 78.6 | — | — | — | — | — | |
| GLaMMMLLMs=LLaVA-7B2026.03 | 78.05 | — | — | 76.6 | — | — | |
| SegLLMMLLMs=LLaVA-7B2026.03 | 77.22 | — | — | 73.41 | — | — | |
| CMIRNetMLLMs=–2026.03 | 77 | — | — | 75.33 | — | — | |
| HeROD-G2026.03 | 76.86 | — | — | — | — | — | |
| AMLTraining Strategy=Standard2026.02 | 75.61 | — | — | 73.19 | — | — | |
| OracleBackbone=ResNet-50, Segmentor=true2024.10 | 75.5 | — | — | — | — | — | |
| SimVG-SegVisual Enc.=BEiT3-B, Text Enc.=BEiT3-B2026.06 | 75.37 | — | — | — | — | — | |
| DETRIS-LTuning Protocol=Parameter Efficient Tuning, Model Scale=Large2025.01 | 75.3 | — | — | — | — | — | |
| UNINEXT-LTraining Data Strategy=Mixed RefCOCO dataset, Model Scale=Large2025.01 | 74.9 | — | — | — | — | — | |
| TALENTPET=true2026.04 | 74.9 | — | — | 72.3 | — | — | |
| PolyFormer-L*Training Data Strategy=Mixed RefCOCO dataset, Model Scale=Large2025.01 | 74.6 | — | — | — | — | — | |
| CARIS*Training Strategy=Standard2026.02 | 74.51 | — | — | 71.86 | — | — | |
| MagNetTraining Strategy=Standard2026.02 | 74.5 | — | — | 71.32 | — | — | |
| DETRIS-BTuning Protocol=Parameter Efficient Tuning, Model Scale=Base2025.01 | 74 | — | — | — | — | — | |
| DETRISPET=true2026.04 | 74 | — | — | — | — | — | |
| PGBDTraining Strategy=Standard2026.02 | 73.84 | — | — | 71.26 | — | — | |
| CGFormerTraining Strategy=Standard2026.02 | 73.76 | — | — | 71 | — | — | |
| RISCLIP-BTuning Protocol=Traditional Full Fine-tuning, Model Scale=Base2025.01 | 73.5 | — | — | — | — | — | |
| RISCLIPPET=true2026.04 | 73.5 | — | — | 70.6 | — | — | |
| DETRIS*Training Strategy=Standard2026.02 | 73.03 | — | — | 70.82 | — | — | |
| LQMFormerBackbone=Swin-B2025.12 | 71.84 | — | — | — | — | — | |
| PixelLM-7BBackbone=CLIP-ViT-L2025.12 | 71.7 | — | — | — | — | — | |
| CoupAlignMLLMs=–2026.03 | 71.56 | — | — | 68.34 | — | — | |
| ETOGPET=true2026.04 | 71.5 | — | — | 68.5 | — | — | |
| MagNetTuning Protocol=Traditional Full Fine-tuning2025.01 | 71.3 | — | — | — | — | — | |
| OmniSegNetBackbone=Swin-B2025.12 | 71.3 | — | — | — | — | — | |
| ReLABackbone=Swin-B2025.12 | 71.02 | — | — | — | — | — | |
| ReLA (2023b)Tuning Protocol=Traditional Full Fine-tuning2025.01 | 71 | — | — | — | — | — | |
| ReLA (2023a)Tuning Protocol=Traditional Full Fine-tuning2025.01 | 71 | — | — | — | — | — | |
| CGFormerTuning Protocol=Traditional Full Fine-tuning2025.01 | 71 | — | — | — | — | — | |
| LAVTTraining Strategy=Standard2026.02 | 70.97 | — | — | 68.38 | — | — | |
| LAVTMLLMs=–2026.03 | 70.97 | — | — | 68.38 | — | — | |
| LAVTVisual Enc.=Swin-B, Text Enc.=BERT-B2026.06 | 70.97 | — | — | — | — | — | |
| LAVT+Backbone=Swin-Base2024.10 | 70.9 | — | — | — | — | — | |
| ReMamberTuning Protocol=Traditional Full Fine-tuning2025.01 | 70.8 | — | — | — | — | — | |
| BarLeRIATuning Protocol=Parameter Efficient Tuning2025.01 | 70.8 | — | — | — | — | — | |
| LISA-7BBackbone=ViT-H2025.12 | 70.8 | — | — | — | — | — | |
| BarLeRIaTraining Strategy=Standard2026.02 | 70.8 | — | — | — | — | — | |
| BarLeRIaPET=true2026.04 | 70.8 | — | — | — | — | — | |
| DMMITuning Protocol=Traditional Full Fine-tuning2025.01 | 69.7 | — | — | — | — | — | |
| GSVA-7BBackbone=ViT-H2025.12 | 69.6 | — | — | — | — | — | |
| ETRISTraining Strategy=Standard2026.02 | 68.51 | — | — | — | — | — | |
| LAVTTuning Protocol=Traditional Full Fine-tuning2025.01 | 68.4 | — | — | — | — | — | |
| LAVTBackbone=Swin-B2025.12 | 68.38 | — | — | — | — | — | |
| CRISTuning Protocol=Traditional Full Fine-tuning2025.01 | 68.1 | — | — | — | — | — | |
| CRISTraining Strategy=Standard2026.02 | 68.08 | — | — | — | — | — | |
| CRISMLLMs=–2026.03 | 68.08 | — | — | — | — | — | |
| CRISVisual Enc.=CLIP-RN101, Text Enc.=CLIP-T2026.06 | 68.08 | — | — | — | — | — | |
| LISA-7BTuning Protocol=Traditional Full Fine-tuning, Model Scale=7B2025.01 | 67.4 | — | — | — | — | — | |
| ETRISTuning Protocol=Parameter Efficient Tuning2025.01 | 66.9 | — | — | — | — | — | |
| ETRISPET=true2026.04 | 66.9 | — | — | — | — | — | |
| DITBackbone=ViT-B2025.12 | 65.52 | — | — | — | — | — | |
| RESTRTuning Protocol=Traditional Full Fine-tuning2025.01 | 60.4 | — | — | — | — | — | |
| LISAMLLMs=LLaVA-7B2026.03 | 59.63 | — | — | 59.43 | — | — | |
| PCNet_SBackbone=ResNet-50, Segmentor=true2024.10 | 56.5 | — | — | — | — | — | |
| B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 54.39 | — | — | — | — | — | |
| B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 54.12 | — | — | — | — | — | |
| B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 54.08 | — | — | — | — | — | |
| B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 54.08 | — | — | — | — | — | |
| B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 53.6 | — | — | — | — | — | |
| B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 53.22 | — | — | — | — | — | |
| B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.89 | — | — | — | — | — | |
| B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.66 | — | — | — | — | — | |
| B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.57 | — | — | — | — | — | |
| B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.5 | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=Step1X-Edit (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 52.37 | — | — | 49.3 | — | — | |
| B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.36 | — | — | — | — | — | |
| B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 52.22 | — | — | — | — | — | |
| IteRPrimEreported_precision=one decimal place2026.03 | 51.6 | — | — | — | — | — | |
| Pseudo-RISVision Backbone=ViT-B, Pretrained Model=SAM, CoCa, CLIP, Training protocol=zero-shot methods w/ additional training2026.05 | 51.42 | — | — | 46.43 | — | — | |
| TAS2026.03 | 50.58 | — | — | — | — | — | |
| TASVision Backbone=ViT-B, Pretrained Model=SAM, BLIP2, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 49.13 | — | — | 38.77 | — | — | |
| HybridGLVision Backbone=ViT-B, Pretrained Model=SAM, CLIP, Training protocol=zero-shot methods w/o additional training2026.05 | 49.13 | — | — | 41.43 | — | — | |
| Global-LocalBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 48.98 | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 47.6 | — | — | — | — | — | |
| VLM-VGVision Backbone=R101, Pretrained Model=COCO∗, VLM-VG∗, Training protocol=zero-shot methods w/ additional training2026.05 | 47.3 | — | — | 40.7 | — | — | |
| RefAMVision Backbone=DiT, Pretrained Model=SAM, FLUX.1-dev (12B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=12B2026.05 | 47.28 | — | — | 42.66 | — | — | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 47.11 | — | — | — | — | — | |
| OursVision Backbone=DiT, Pretrained Model=FLUX.2-klein (9B), Training protocol=zero-shot methods w/o additional training, Number of Parameters=9B2026.05 | 46.72 | — | — | 43.25 | — | — | |
| Global-LocalBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 46.17 | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 45.96 | — | — | — | — | — | |
| B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 45.94 | — | — | — | — | — | |
| PPTBackbone=ViT-Base, Segmentor=true2024.10 | 45.8 | — | — | — | — | — | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 44.74 | — | — | — | — | — | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 44.44 | — | — | — | — | — | |
| CLIPBackbone=ResNet-50, Segmentor=true2024.10 | 42.7 | — | — | — | — | — | |
| B2GBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Zero-shot Evaluation Protocol=true, CT (Confidence Thresholding)=true, SG (Spatial Guiding)=true2025.09 | 42.12 | — | — | — | — | — |