Referring Image Segmentation on OmniRef Text (test)
65.3cIoUGSVA-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GSVA-7BBackbone=ViT-H, Text Encoder=Vicuna-7B2025.12 | 65.3 | 67.57 | 63.44 | |
| LISA-7BBackbone=ViT-H, Text Encoder=Vicuna-7B2025.12 | 64.95 | 66.02 | — | |
| OmniSegNetBackbone=Swin-B, Text Encoder=BERT2025.12 | 64.92 | 66.44 | 62.56 | |
| ReLABackbone=Swin-B, Text Encoder=BERT2025.12 | 63.4 | 64.75 | 57.97 | |
| VRP-SAM+ReLABackbone=ViT-L, Text Encoder=BERT2025.12 | 63.4 | 64.75 | 57.97 | |
| DCAMA+ReLABackbone=Swin-B, Text Encoder=BERT2025.12 | 63.4 | 64.75 | 57.97 | |
| DITBackbone=ViT-B, Text Encoder=BERT2025.12 | 60.72 | 62.73 | — |