Generalized Referring Expression Segmentation on gRefCOCO (testB)
73.1cIoUPSALM-G5 (ft)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| PSALM-G5 (ft)Category=VLM based, Fine-tuning=true, Training Data=Ground-V2025.05 | 73.1 | 78.9 | 90.6 | — | — | |
| PSALM-G5Category=VLM based, Fine-tuning=false, Training Data=Ground-V2025.05 | 72.3 | 72.7 | 84.6 | — | — | |
| PSALM (ft)Category=VLM based, Fine-tuning=true2025.05 | 71.3 | — | 79.6 | — | — | |
| DeRIS-LVisual Encoder=Swin-B, Textual Encoder=BEIT3-L2025.07 | 67.38 | 67.99 | 66.81 | — | — | |
| HiMTok-8BZero-Shot=false2026.03 | 65.8 | 64.1 | — | — | — | |
| GETok-RLTraining M-Dec.=false2025.12 | 65.6 | — | — | — | — | |
| LISA-G5 (ft)Category=VLM based, Fine-tuning=true, Training Data=Ground-V2025.05 | 65 | 66.3 | 63.2 | — | — | |
| DeRIS-BVisual Encoder=Swin-S, Textual Encoder=BEIT3-B2025.07 | 64.65 | 65.63 | 63.44 | — | — | |
| AnchorSeg-LLaVA1.5-7BBackbone=LLaVA1.5-7B, fine-tuned=true2026.04 | 64.56 | 68.25 | 68.27 | — | — | |
| GETok-SFTTraining M-Dec.=false2025.12 | 64.1 | — | — | — | — | |
| InstAlignBackbone=Swin-B2024.11 | 63.88 | 65.74 | 70.72 | — | — | |
| GSVATraining M-Dec.=true2025.12 | 63.8 | — | — | — | — | |
| SAM4MLLM-8BVisual Encoder=SAM-EfficientViT-XL1, Textual Encoder=Qwen-VL-7B2025.07 | 63.42 | 65.29 | 59.99 | — | — | |
| SAM4MLLM-8BBackbone=LLaVA1.6, External datasets=true2024.11 | 63.42 | 65.29 | 59.99 | — | — | |
| SAM4MLLM-8BZero-Shot=false2026.03 | 63.4 | 65.3 | — | — | — | |
| SAM4MLLM-7BType=MLLM-based Segmentation Network, SAM-based=false, Evaluation Protocol=Supervised2024.11 | 63.2 | — | — | — | — | |
| GSVA (ft)Category=VLM based, Fine-tuning=true2025.05 | 63.2 | 62.2 | 62.5 | — | — | |
| EEVGVisual Encoder=ViT-B, Textual Encoder=BERT-B2025.07 | 62.77 | 62.79 | — | — | — | |
| MABPBackbone=Swin-B2024.11 | 62.76 | 64.04 | — | — | — | |
| MABPVisual Encoder=Swin-B, Textual Encoder=BERT2024.05 | 62.75 | 64.01 | — | — | — | |
| CoHDBackbone=Swin-B2024.05 | 62.63 | 63.6 | 60.37 | — | — | |
| CoHDBackbone=Swin-B2024.11 | 62.63 | 63.6 | 60.37 | — | — | |
| SAM4MLLM-7BBackbone=SAM-EffViT-XL1, External datasets=true2024.11 | 62.35 | 63.71 | 61.25 | — | — | |
| UGround-LLaVA1.5-7Bfine-tuned=true2025.10 | 61.84 | 66.85 | 71.22 | — | — | |
| LISATraining M-Dec.=true2025.12 | 61.8 | — | — | — | — | |
| VisHarnessTraining Data Ratio=0.7%2026.05 | 61.35 | 62.44 | — | — | 63.58 | |
| ReLATraining M-Dec.=true2025.12 | 61 | — | — | — | — | |
| ReLAVisual Encoder=Swin-B, Textual Encoder=BERT2024.05 | 60.97 | 61.05 | — | — | — | |
| HieA2GVisual Encoder=Swin-B, Textual Encoder=RoBERTa-B2025.07 | 60.8 | 62.8 | 61 | — | — | |
| LISA-7BVisual Encoder=SAM-ViT-H, Textual Encoder=Vicuna-7B2025.07 | 60.63 | 58.84 | 51.91 | — | — | |
| LISA-V-7B (ft)Backbone=SAM-ViT-H, Finetuned=true2024.05 | 60.63 | 58.84 | 51.91 | — | — | |
| LISA-Vicuna-7Bfine-tuned=true2025.10 | 60.63 | 58.84 | 51.91 | — | — | |
| LISA-7BBackbone=SAM-ViT-H, External datasets=true2024.11 | 60.63 | 58.84 | 51.91 | — | — | |
| LISA-Vicuna-7BBackbone=Vicuna-7B, fine-tuned=true2026.04 | 60.63 | 58.84 | 51.91 | — | — | |
| LISA(FT)Training Data Ratio=100%2026.05 | 60.63 | 58.84 | — | — | 62.9 | |
| LISA-7B(ft)Zero-Shot=false2026.03 | 60.6 | 58.8 | — | — | — | |
| GSVA(ft)LLM Type=Vicuna-7B, Zero-Shot=false, Fine-tuning=true2024.03 | 60.5 | 62.2 | — | — | — | |
| GSVA-7B(ft)Zero-Shot=false2026.03 | 60.5 | 62.2 | — | — | — | |
| GSVA-7BVisual Encoder=SAM-ViT-H, Textual Encoder=Vicuna-7B2025.07 | 60.47 | 62.23 | 60.56 | — | — | |
| GSVA-V-7B (ft)Backbone=SAM-ViT-H, Finetuned=true2024.05 | 60.47 | 62.23 | 60.56 | — | — | |
| GSVA-Vicuna-7Bfine-tuned=true2025.10 | 60.47 | 62.23 | 60.56 | — | — | |
| GSVA-7BBackbone=SAM-ViT-H, External datasets=true2024.11 | 60.47 | 62.23 | 60.56 | — | — | |
| GSVA-Vicuna-7BBackbone=Vicuna-7B, fine-tuned=true2026.04 | 60.47 | 62.23 | 60.56 | — | — | |
| GSVA(FT)Training Data Ratio=100%2026.05 | 60.47 | 62.23 | — | — | 65.6 | |
| CoHDBackbone=Swin-T2024.05 | 60.32 | 61.78 | 60 | — | — | |
| LISA(ft)LLM Type=Vicuna-7B, Zero-Shot=false, Fine-tuning=true2024.03 | 60.3 | 61.3 | — | — | — | |
| GSVALLM Type=Vicuna-7B, Zero-Shot=false, Fine-tuning=false2024.03 | 60.3 | 61.3 | — | — | — | |
| GSVA-7BType=MLLM-based Segmentation Network, SAM-based=true, Evaluation Protocol=Supervised2024.11 | 60.3 | — | — | — | — | |
| GSVACategory=VLM based, Fine-tuning=false2025.05 | 60.3 | 61.3 | 60.6 | — | — | |
| LISA (ft)Category=VLM based, Fine-tuning=true2025.05 | 60.3 | 61.3 | 58.4 | — | — | |
| GSVA-7BZero-Shot=false2026.03 | 60.3 | 61.3 | — | — | — | |
| GSVA-V-7BBackbone=SAM-ViT-H, Finetuned=false2024.05 | 60.26 | 61.34 | 58.42 | — | — | |
| GSVA-Vicuna-7Bfine-tuned=false2025.10 | 60.26 | 61.34 | 58.42 | — | — | |
| GSVA-Vicuna-7BBackbone=Vicuna-7B2026.04 | 60.26 | 61.34 | 58.42 | — | — | |
| VisHarness-SFTTraining Data Ratio=0.7%2026.05 | 60.25 | 61.72 | — | — | 62.32 | |
| CGFormerVisual Encoder=Swin-B, Textual Encoder=BERT2024.05 | 60.18 | 61.09 | — | — | — | |
| ReLA2026.01 | 60.15 | 61.29 | — | — | — | |
| DMMI∗Backbone=Swin-B2024.05 | 60.01 | 60.09 | 54.58 | — | — | |
| ReLAZero-Shot=false2024.03 | 59.9 | 61 | — | — | — | |
| ReLACategory=Specialist, Fine-tuning=false2025.05 | 59.9 | 61 | 58.4 | — | — | |
| ReLAVisual Encoder=Swin-B, Textual Encoder=BERT-B2025.07 | 59.88 | 61.02 | 58.4 | — | — | |
| ReLA2023.06 | 59.88 | 61.02 | — | — | — | |
| ReLABackbone=Swin-B2024.05 | 59.88 | 61.02 | 58.4 | — | — | |
| ReLA2025.10 | 59.88 | 61.02 | 58.4 | — | — | |
| ReLABackbone=Swin-B2024.11 | 59.88 | 61.02 | 58.4 | — | — | |
| ReLA2026.04 | 59.88 | 61.02 | 58.4 | — | — | |
| ReLATraining Data Ratio=100%2026.05 | 59.88 | 61.02 | — | — | 64.4 | |
| LAVT+RELADecoder=ReLA2023.06 | 58.24 | 59.83 | — | — | — | |
| LAVT+ReLAbase_method=LAVT2026.01 | 58.24 | 59.83 | — | — | — | |
| ReLA†Backbone=Swin-T2024.05 | 56.86 | 57.37 | 50.31 | — | — | |
| VLT+RELADecoder=ReLA2023.06 | 56.22 | 57.36 | — | — | — | |
| VLT+ReLAbase_method=VLT2026.01 | 56.22 | 57.36 | — | — | — | |
| LAVTTraining M-Dec.=true2025.12 | 55.8 | — | — | — | — | |
| LAVTVisual Encoder=Swin-B, Textual Encoder=BERT-B2025.07 | 55.04 | 55.83 | 48.46 | — | — | |
| LAVT2023.06 | 55.04 | 55.83 | — | — | — | |
| LAVTBackbone=Swin-B2024.05 | 55.04 | 55.83 | 48.46 | — | — | |
| LAVTVisual Encoder=Swin-B, Textual Encoder=BERT2024.05 | 55.04 | 55.83 | — | — | — | |
| LAVT2026.01 | 55.04 | 55.83 | — | — | — | |
| LAVT2025.10 | 55.04 | 55.83 | 48.46 | — | — | |
| LAVTBackbone=Swin-B2024.11 | 55.04 | 55.83 | 48.46 | — | — | |
| LAVT2026.04 | 55.04 | 55.83 | 48.46 | — | — | |
| LAVTTraining Data Ratio=100%2026.05 | 55.04 | 55.83 | — | — | 59.7 | |
| LAVTType=Segmentation Specialist, SAM-based=false, Evaluation Protocol=Supervised2024.11 | 55 | — | — | — | — | |
| LAVTCategory=Specialist, Fine-tuning=false2025.05 | 55 | 55.8 | 48.5 | — | — | |
| LISA-G5Category=VLM based, Fine-tuning=false, Training Data=Ground-V2025.05 | 53.9 | 51.3 | 44.1 | — | — | |
| HyperSegType=MLLM-based Segmentation Network, SAM-based=false, Evaluation Protocol=Zero-shot2024.11 | 52.5 | — | — | — | — | |
| CRISVisual Encoder=CLIP-ResNet-101, Textual Encoder=CLIP-B2025.07 | 51.04 | 51.79 | — | — | — | |
| CRIS2023.06 | 51.04 | 51.79 | — | — | — | |
| CRISBackbone=CLIP-R1012024.05 | 51.04 | 51.79 | — | — | — | |
| CRISVisual Encoder=CLIP-R 101, Textual Encoder=CLIP2024.05 | 51.04 | 51.79 | — | — | — | |
| CRIS2026.01 | 51.04 | 51.79 | — | — | — | |
| CRIS2025.10 | 51.04 | 51.79 | — | — | — | |
| CRISBackbone=ResNet-1012024.11 | 51.04 | 51.79 | — | — | — | |
| CRIS2026.04 | 51.04 | 51.79 | — | — | — | |
| CRISTraining Data Ratio=100%2026.05 | 51.04 | 51.79 | — | — | 56.95 | |
| CRISType=Segmentation Specialist, SAM-based=false, Evaluation Protocol=Supervised2024.11 | 51 | — | — | — | — | |
| SELF1E-8BZero-Shot=true2026.03 | 50.9 | 45.6 | — | — | — | |
| PSALMLLM Type=Phi-1.5 (1.3B), Zero-Shot=true2024.03 | 50.6 | 52.5 | — | — | — | |
| PSALMType=MLLM-based Segmentation Network, SAM-based=false, Evaluation Protocol=Zero-shot2024.11 | 50.6 | — | — | — | — | |
| PSALMCategory=VLM based, Fine-tuning=false2025.05 | 50.6 | 52.5 | 25.6 | — | — |