Semantic Segmentation on SUIM
62mIoUMPA (w Source-Training)
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| MPA (w Source-Training)Shot=5-shot, Source-Training=yes2026.02 | 62 | — | — | — | — | — | — | — | — | — | |
| INSID3Encoder=DINOv3, #Param=304 M, Training Protocol=Training free: Unsupervised pre-training, Shot count=52026.03 | 61.7 | — | — | — | — | — | — | — | — | — | |
| MPA (w/o Source-Training)Shot=5-shot, Source-Training=no2026.02 | 61.1 | — | — | — | — | — | — | — | — | — | |
| GF-SAM† + our debiasEncoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 59.2 | — | — | — | — | — | — | — | — | — | |
| GF-SAM†Encoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 58.6 | — | — | — | — | — | — | — | — | — | |
| GF-SAMEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 58.1 | — | — | — | — | — | — | — | — | — | |
| MPA (w Source-Training)Shot=1-shot, Source-Training=yes2026.02 | 55.5 | — | — | — | — | — | — | — | — | — | |
| INSID3Encoder=DINOv3, #Param=304 M, Training Protocol=Training free, Supervision Type=Unsupervised pre-training2026.03 | 54.9 | — | — | — | — | — | — | — | — | — | |
| SINEEncoder=DINOv2, #Param=373 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 54.8 | — | — | — | — | — | — | — | — | — | |
| MPA (w/o Source-Training)Shot=1-shot, Source-Training=no2026.02 | 54.2 | — | — | — | — | — | — | — | — | — | |
| GF-SAMEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 53.1 | — | — | — | — | — | — | — | — | — | |
| SegIC (COCO)Encoder=DINOv2, #Param=310 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision, dataset split=COCO2026.03 | 52.9 | — | — | — | — | — | — | — | — | — | |
| GF-SAM† + debiasEncoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training, Debias=True2026.03 | 52.9 | — | — | — | — | — | — | — | — | — | |
| SegICEncoder=DINOv2, #Param=310 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 52.5 | — | — | — | — | — | — | — | — | — | |
| SINEEncoder=DINOv2, #Param=373 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 50.7 | — | — | — | — | — | — | — | — | — | |
| MatcherEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 50.6 | — | — | — | — | — | — | — | — | — | |
| GF-SAM†Encoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 50.5 | — | — | — | — | — | — | — | — | — | |
| DiffewSEncoder=Stable Diffusion, #Param=890 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 49.8 | — | — | — | — | — | — | — | — | — | |
| DiffewSEncoder=Stable Diffusion, #Param=890 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 48.9 | — | — | — | — | — | — | — | — | — | |
| MatcherEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 44.1 | — | — | — | — | — | — | — | — | — | |
| ABCDFSSShot=5-shot2026.02 | 41.3 | — | — | — | — | — | — | — | — | — | |
| PATNetShot=5-shot2026.02 | 40.2 | — | — | — | — | — | — | — | — | — | |
| ABCDFSSShot=1-shot2026.02 | 35.1 | — | — | — | — | — | — | — | — | — | |
| SegGPTEncoder=ViT, #Param=354 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 34.9 | — | — | — | — | — | — | — | — | — | |
| PxMtchShot=1-shot2026.02 | 34.8 | — | — | — | — | — | — | — | — | — | |
| RemDiffShot=1-shot2026.02 | 34.7 | — | — | — | — | — | — | — | — | — | |
| SegGPTEncoder=ViT, #Param=354 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 33.7 | — | — | — | — | — | — | — | — | — | |
| PATNetShot=1-shot2026.02 | 32.1 | — | — | — | — | — | — | — | — | — | |
| SCLShot=1-shot2026.02 | 31.8 | — | — | — | — | — | — | — | — | — | |
| HDMNetShot=5-shot2026.02 | 30.9 | — | — | — | — | — | — | — | — | — | |
| HSNetShot=1-shot2026.02 | 28.8 | — | — | — | — | — | — | — | — | — | |
| PerSAMEncoder=SAM, #Param=640 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 28.7 | — | — | — | — | — | — | — | — | — | |
| RestNetShot=1-shot2026.02 | 25.2 | — | — | — | — | — | — | — | — | — | |
| HDMNetShot=1-shot2026.02 | 23.4 | — | — | — | — | — | — | — | — | — | |
| BaselineSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 60.46 | 84.57 | 66.86 | 21.07 | 56.72 | 60.92 | 60.49 | 65.81 | 67.23 | |
| DW-NetSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 61.95 | 85.41 | 66.88 | 15.44 | 66.94 | 62.64 | 62.88 | 68.84 | 66.56 | |
| GUPDMSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 62.67 | 85.72 | 65.04 | 19.72 | 67.27 | 63.12 | 65.57 | 66.93 | 68.03 | |
| HCLR-NetSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 61.25 | 84.92 | 66.69 | 23.44 | 60.61 | 61.37 | 63.09 | 69.66 | 60.17 | |
| SDAR-NetSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 63.58 | 85.69 | 73.03 | 18.22 | 65.59 | 70.52 | 57.58 | 73.37 | 64.64 | |
| Semi-UIRSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 61.98 | 83.54 | 71.52 | 17.62 | 64.59 | 65.07 | 63.22 | 71.1 | 59.2 | |
| U-shapeSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 58.31 | 85.84 | 71.3 | 5.78 | 62.75 | 46.54 | 60.17 | 69.84 | 64.28 | |
| UcolorSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 57.8 | 84.01 | 66.57 | 1.96 | 62.05 | 61.07 | 61.21 | 64.28 | 61.22 | |
| UDCPSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 59.03 | 85.12 | 68 | 7.16 | 56.86 | 63.84 | 58.76 | 70.66 | 61.82 | |
| UWCNNSegmentation Backbone=U-Net, Evaluation Protocol=Fine-tuned on enhanced images2026.04 | — | 60.51 | 85.4 | 70.23 | 20.73 | 61.42 | 56.46 | 60.72 | 64.58 | 64.57 |