Referring Remote Sensing Image Segmentation on RRSIS-D (val)
68.1mIoU (Mean IoU)Qwen3-VL-SAM
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-SAMLLM=Qwen3-VL-2B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | 68.1 | — | — | — | — | — | — | — | |
| GeoPixelLLM=InternLM2-7B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 68 | — | — | — | — | — | — | — | |
| SegEarth-R1LLM=phi-1.5-1.3B, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 67.6 | — | — | — | — | — | — | — | |
| BTDNetLLM=BERT-base, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 66.9 | — | — | — | — | — | — | — | |
| SBANetLLM=BERT-base, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 66.7 | — | — | — | — | — | — | — | |
| ICIPNetVisual Encoder=Swin-B, Text Encoder=BERT2026.05 | 65.19 | 75.17 | 67.83 | 56.6 | 44.06 | 24.99 | 77.69 | — | |
| RMSINVisual Encoder=Swin-B, Text Encoder=BERT2023.12 | 65.1 | 74.66 | 68.22 | 57.41 | 45.29 | 24.43 | 78.27 | — | |
| RMSINLLM=BERT-base, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 65.1 | — | — | — | — | — | — | — | |
| RemoteSAMVisual Encoder=Swin-B, Text Encoder=BERT2026.05 | 64.11 | 73.6 | 66.59 | 57.47 | 43.77 | 25.75 | 77.04 | — | |
| MAFNVisual Encoder=Swin-B, Text Encoder=BERT2026.05 | 64.08 | 74.63 | 67.22 | 56.47 | 43.43 | 24.75 | 77.23 | — | |
| FIANetVisual Encoder=Swin-B, Text Encoder=BERT2026.05 | 63.87 | 73.74 | 67.64 | 55.92 | 43.62 | 25.23 | 76.95 | — | |
| DiffRISLLM=CLIP, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 63.6 | — | — | — | — | — | — | — | |
| RMSINVisual Encoder=Swin-B, Text Encoder=BERT2026.05 | 63.14 | 72.7 | 64.89 | 54.43 | 42.3 | 22.82 | 76.31 | — | |
| LAVTLLM=BERT-base, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 61.5 | — | — | — | — | — | — | — | |
| LAVTVisual Encoder=Swin-B, Text Encoder=BERT2023.12 | 61.46 | 69.54 | 63.51 | 53.16 | 43.97 | 24.25 | 77.59 | — | |
| RRSISLLM=BERT-base, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 60.2 | — | — | — | — | — | — | — | |
| LGCEVisual Encoder=Swin-B, Text Encoder=BERT2023.12 | 60.16 | 68.1 | 60.52 | 52.24 | 42.24 | 23.85 | 76.68 | — | |
| CMPC+Visual Encoder=R-101, Text Encoder=LSTM2023.12 | 51.41 | 59.19 | 49.36 | 38.67 | 25.91 | 8.16 | 70.14 | — | |
| BRINetVisual Encoder=R-101, Text Encoder=LSTM2023.12 | 51.14 | 58.79 | 49.54 | 39.65 | 28.21 | 9.19 | 70.73 | — | |
| CMPCVisual Encoder=R-101, Text Encoder=LSTM2023.12 | 50.41 | 57.93 | 48.85 | 38.5 | 25.28 | 9.31 | 70.15 | — | |
| LSCMVisual Encoder=R-101, Text Encoder=LSTM2023.12 | 50.36 | 57.12 | 48.04 | 37.87 | 26.37 | 7.93 | 69.28 | — | |
| CSMAVisual Encoder=R-101, Text Encoder=None2023.12 | 48.85 | 55.68 | 48.04 | 38.27 | 26.55 | 9.02 | 69.68 | — | |
| RRNVisual Encoder=R-101, Text Encoder=LSTM2023.12 | 46.06 | 51.09 | 42.47 | 33.04 | 20.8 | 6.14 | 66.53 | — | |
| PixelLMLLM=Vicuna-7B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 33.9 | — | — | — | — | — | — | — | |
| LISALLM=Vicuna-7B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 27.8 | — | — | — | — | — | — | — | |
| NExT-ChatLLM=Vicuna-7B, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | 27 | — | — | — | — | — | — | — | |
| GPT-5-SAMLLM=GPT-5, Trained on RS data: LLM=false, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | 25.8 | — | — | — | — | — | — | — | |
| GPT-Image-1LLM=GPT-5, Trained on RS data: LLM=false, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | 20.1 | — | — | — | — | — | — | — | |
| CMPC+Pub=TPAMI'21, Method Category=Segmentation Specialists2025.12 | — | — | — | — | — | — | — | 51.4 | |
| GeoGroundPub=arXiv'24, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 61.1 | |
| GeoPixelPub=ICML'25, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 68 | |
| LAVTPub=CVPR'22, Method Category=Segmentation Specialists2025.12 | — | — | — | — | — | — | — | 61.5 | |
| LISAPub=CVPR'24, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 27.8 | |
| PixelLMPub=CVPR'24, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 33.9 | |
| RIS-DMMIPub=CVPR'23, Method Category=Segmentation Specialists2025.12 | — | — | — | — | — | — | — | 60.7 | |
| RMSINPub=CVPR'24, Method Category=Segmentation Specialists2025.12 | — | — | — | — | — | — | — | 65.1 | |
| SegEarth-R1Pub=arXiv'25, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 67.6 | |
| SegEarth-R2Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 68.8 | |
| Text4Seg++Pub=arXiv'25, Method Category=MLLM based segmentation2025.12 | — | — | — | — | — | — | — | 64.1 |