Referring Image Segmentation on RefCOCO RefCOCO+ RefCOCOg Overall Average
42.33oIoU+B2Gfix
Evaluation Results
| Method | Links | |
|---|---|---|
| +B2GfixBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.33 | |
| +B2Gdyn-clusBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.31 | |
| +B2Gdyn-permBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 42.02 | |
| +B2GfixBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.85 | |
| +B2Gdyn-permBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.75 | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.58 | |
| +B2Gdyn-clusBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.54 | |
| +B2Gdyn-permBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.24 | |
| +B2GfixBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 41.05 | |
| Global-LocalBackbone=SigLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 40.69 | |
| +B2GfixBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.78 | |
| +B2Gdyn-permBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.54 | |
| +B2GBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 39.5 | |
| +B2Gdyn-clusBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 39.38 | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 38.68 | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 38.58 | |
| HybridGLBackbone=CLIP ViT-B/16, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 38.02 | |
| Global-LocalBackbone=DFN VIT-H/14, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 36.74 | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 36.73 | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=Mask2Former, Evaluation Protocol=Zero-shot2025.09 | 36.55 | |
| +B2GBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 35.23 | |
| Global-LocalBackbone=CLIP ViT-B/16, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 35.19 | |
| HybridGLBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 34.9 | |
| Global-LocalBackbone=CLIP ViT-B/32, Pre-trained Segmentor=SAM, Evaluation Protocol=Zero-shot2025.09 | 31.55 |