Reasoning Segmentation on EarthReason (val)
76.35gIoUGRTO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GRTO2026.05 | 76.35 | 78.32 | — | — | |
| B-GRTO2026.05 | 76.25 | 77.59 | — | — | |
| GRPO2026.05 | 73.76 | 73.32 | — | — | |
| UniGeoSegdomain-specific=true2026.05 | 72.54 | 73.3 | — | — | |
| Think2Seg-RSModel=Qwen-2.5-VL-7B + SAM2-Small2025.12 | 72.46 | 73.62 | — | — | |
| SegEarth-R2-3BLLM=Phi-2-2.7B2025.12 | 72.3 | 68.1 | — | — | |
| Think2Seg-RS2026.05 | 72.16 | 72.87 | — | — | |
| Text4Seg++LLM=Qwen2-7B2025.12 | 71.9 | 69.8 | — | — | |
| Think2Seg-RSModel=Qwen-2.5-VL-3B + SAM2-Small2025.12 | 69.3 | 71.28 | — | — | |
| RemoteReasoner2026.04 | 69.29 | — | — | 67.04 | |
| RemoteReasonerdomain-specific=true2026.05 | 69.13 | 67.8 | — | — | |
| RemoteReasonerModel=Qwen-2.5-VL-7B + SAM22025.12 | 69.02 | 67.8 | — | — | |
| RemoteReasonerLLM=Qwen2.5-7B2025.12 | 69 | 67.8 | — | — | |
| SegEarth-R1LLM=Phi-1.5-1.3B2025.12 | 68.6 | 64.1 | — | — | |
| SegEarth-R1Model=Swin-B (88M) + Phi-1.5-1.3B2025.12 | 68.6 | 64.13 | — | — | |
| SegEarth-R1Evaluation Protocol=Full Amount Fine-tune2025.09 | 68.6 | — | — | — | |
| SegEarth-R1domain-specific=true2026.05 | 68.25 | 64.13 | — | — | |
| Seg-Zero2026.05 | 67.15 | 65.41 | — | — | |
| PSALMModel=Swin-B (88M) + Phi-1.5-1.3B2025.12 | 66.61 | 62.03 | — | — | |
| PSALMEvaluation Protocol=Full Amount Fine-tune2025.09 | 66.61 | — | — | — | |
| PSALM2026.05 | 66.61 | 62.03 | — | — | |
| VisionReasoner2026.05 | 66.58 | 67.13 | — | — | |
| VisionReasonerModel=Qwen-2.5-VL-7B + SAM2-Large2025.12 | 66.5 | 67.98 | — | — | |
| Seg-ZeroModel=Qwen-2.5-VL-7B + SAM2-Large2025.12 | 63 | 61.93 | — | — | |
| LISAModel=CLIP-L (304M) + Vicuna-7B2025.12 | 61.04 | 57.39 | — | — | |
| LISAEvaluation Protocol=Full Amount Fine-tune2025.09 | 61.04 | — | — | — | |
| LISALLM=Vicuna-7B2025.12 | 61 | 57.4 | — | — | |
| LISA2026.05 | 59.1 | 57.39 | — | — | |
| PixelLMModel=CLIP-L (304M) + Vicuna-7B2025.12 | 57.94 | 57.79 | — | — | |
| PixelLMEvaluation Protocol=Full Amount Fine-tune2025.09 | 57.94 | — | — | — | |
| PixelLM2026.05 | 57.94 | 57.79 | — | — | |
| PixelLMLLM=Vicuna-7B2025.12 | 57.9 | 57.8 | — | — | |
| Geo-R1 (GRPO)Evaluation Protocol=10-shot Fine-tune, Number of training samples=2402025.09 | 57.78 | — | — | — | |
| SegEarth-R1Evaluation Protocol=10-shot Fine-tune, Number of training samples=2402025.09 | 56.4 | — | — | — | |
| Geo-R1 (DAPO)Evaluation Protocol=10-shot Fine-tune, Number of training samples=2402025.09 | 55.56 | — | — | — | |
| Geo-R1 (GRPO)Evaluation Protocol=5-shot Fine-tune, Number of training samples=1202025.09 | 54.73 | — | — | — | |
| Geo-R1 (DAPO)Evaluation Protocol=5-shot Fine-tune, Number of training samples=1202025.09 | 54.46 | — | — | — | |
| RemoteAgent2026.04 | 52.22 | — | — | 55.6 | |
| GeoPixeldomain-specific=true2026.05 | 52.13 | 54.23 | — | — | |
| Geo-R1 (GRPO)Evaluation Protocol=1-shot Fine-tune, Number of training samples=242025.09 | 50.3 | — | — | — | |
| Geo-R1 (DAPO)Evaluation Protocol=1-shot Fine-tune, Number of training samples=242025.09 | 50.09 | — | — | — | |
| SegEarth-R1Evaluation Protocol=5-shot Fine-tune, Number of training samples=1202025.09 | 45.37 | — | — | — | |
| SegEarth-R1Evaluation Protocol=1-shot Fine-tune, Number of training samples=242025.09 | 42.47 | — | — | — | |
| Qwen2.5-VL-7B2026.04 | 41.8 | — | — | 38.77 | |
| Qwen2.5-VL w/ thinkingEvaluation Protocol=Zero-shot Baseline2025.09 | 19.35 | — | — | — | |
| DeepSeek-VL2-tiny2026.04 | 18.62 | — | — | 17.51 | |
| GeoChat2026.04 | 11.44 | — | — | 12.57 | |
| GPT-5-SAMLLM=GPT-5, Trained on RS data: LLM=false, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | — | — | 46 | — | |
| GPT-Image-1LLM=GPT-5, Trained on RS data: LLM=false, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | — | — | 38.4 | — | |
| LISALLM=Vicuna-7B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | — | — | 61 | — | |
| PixelLMLLM=Vicuna-7B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | — | — | 57.9 | — | |
| PSALMLLM=phi-1.5-1.3B, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | — | — | 66.6 | — | |
| Qwen3-VL-SAMLLM=Qwen3-VL-2B, Trained on RS data: LLM=LoRA, Trained on RS data: Mask Decoder=false, Trained on RS data: Extra=false2026.02 | — | — | 70.6 | — | |
| SegEarth-R1LLM=phi-1.5-1.3B, Trained on RS data: LLM=true, Trained on RS data: Mask Decoder=true, Trained on RS data: Extra=true2026.02 | — | — | 68.6 | — |