Semantic Segmentation on PASCAL VOC (val)
4,420mIoUMaskContrast
Evaluation Results
| Method | Links | |||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MaskContrastInitialization=IN Sup. Init., Saliency Mask=Supervised, Protocol=K-Means2021.02 | 4,420 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskContrastInitialization=IN Sup. Init., Saliency Mask=Unsupervised, Protocol=K-Means2021.02 | 4,160 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskContrastInitialization=MoCo Init., Saliency Mask=Supervised, Protocol=K-Means2021.02 | 3,890 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskContrastInitialization=MoCo Init., Saliency Mask=Unsupervised, Protocol=K-Means2021.02 | 3,500 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IICCategory=Clustering based, Protocol=K-Means2021.02 | 980 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ColorizationCategory=Proxy task based, Protocol=K-Means2021.02 | 490 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ImageNet (IN) Classifier (Supervised)Protocol=K-Means2021.02 | 470 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Inst. Discr.Category=Contrastive learning based, Protocol=K-Means2021.02 | 440 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SWAVCategory=Contrastive learning based, Protocol=K-Means2021.02 | 440 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CMPCategory=Proxy task based, Protocol=K-Means2021.02 | 430 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MoCo v2Category=Contrastive learning based, Protocol=K-Means2021.02 | 430 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Co-OccurenceCategory=Proxy task based, Protocol=K-Means2021.02 | 400 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InfoMinCategory=Contrastive learning based, Protocol=K-Means2021.02 | 370 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SCANVL-Model=CLIP VIT-L/14, Training Dataset=COCO2023.12 | 97.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SCANVL-Model=CLIP VIT-B/16, Training Dataset=COCO2023.12 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EBSegVLM=CLIP ViT-L/14, Training Dataset=COCO-Stuff2024.06 | 96.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EBSegVLM=CLIP ViT-B/16, Training Dataset=COCO-Stuff2024.06 | 94.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SANVLM=CLIP ViT-L/14, Training Dataset=COCO-Stuff2024.06 | 94.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SANVL-Model=CLIP VIT-L/14, Training Dataset=COCO2023.12 | 94.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OVSegVLM=CLIP ViT-L/14, Training Dataset=COCO-Stuff2024.06 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OVSegVL-Model=CLIP VIT-L/14, Training Dataset=COCO2023.12 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SANVLM=CLIP ViT-B/16, Training Dataset=COCO-Stuff2024.06 | 94 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SANVL-Model=CLIP VIT-B/16, Training Dataset=COCO2023.12 | 94 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OVSegVLM=CLIP ViT-B/16, Training Dataset=COCO-Stuff+COCO Caption[4]2024.06 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OVSegVL-Model=CLIP VIT-B/16, Training Dataset=COCO2023.12 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimSegVLM=CLIP ViT-L/14, Training Dataset=COCO-Stuff2024.06 | 92.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimSeg†VL-Model=CLIP VIT-L/14, Training Dataset=COCO2023.12 | 92.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimSegVLM=CLIP ViT-B/16, Training Dataset=COCO-Stuff2024.06 | 91.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimSeg†VL-Model=CLIP VIT-B/16, Training Dataset=COCO2023.12 | 91.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MAFTVL-Model=CLIP ViT-B/16, Training Dataset=COCO2023.12 | 90 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ZegFormerVLM=CLIP ViT-B/16, Training Dataset=COCO-Stuff2024.06 | 89.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SimSegVL-Model=CLIP VIT-B/16, Training Dataset=COCO2023.12 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv3Arch=7B/16, Probe Type=Linear Probing, Backbone Status=Frozen, Resolution=560x5602026.02 | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VersaViTArch=H/14, Probe Type=Linear Probing, Backbone Status=Frozen, Resolution=560x5602026.02 | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ModuSegType=M, Time (min)=84, GPU (G)=5.32026.04 | 86.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mask2FormerBackbone=Swin-L, Supervision=M2024.03 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReCLIP++Publication=Ours, Setting=C-USS2024.08 | 85.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UPLiFTParams (M)=0.8, Time (ms)=79.4, Backbone=DINOv2-S/142026.01 | 85.21 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.51 | |
| LoftUpParams (M)=4.3, Time (ms)=223.5, Backbone=DINOv2-S/142026.01 | 84.63 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.33 | |
| JAFARParams (M)=0.7, Time (ms)=111.7, Backbone=DINOv2-S/142026.01 | 84.38 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.22 | |
| AnyUpParams (M)=0.9, Time (ms)=146.7, Backbone=DINOv2-S/142026.01 | 84.33 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.23 | |
| FeatUpParams (M)=0.2, Time (ms)=109.6, Backbone=DINOv2-S/142026.01 | 83.52 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.06 | |
| ODISEPre-training Dataset=COCO, Pre-training Supervision=mask2026.03 | 82.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FMA-WSSSBackbone=Swin-L, Supervision=I+C+S2024.03 | 82.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LiFT-2×Params (M)=1.2, Time (ms)=3.8, Backbone=DINOv2-S/142026.01 | 82.46 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.73 | |
| DeepLabV2Backbone=ViT-B/16, Supervision=Fully-supervised2024.02 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DHRBackbone=Swin-L, Supervision=I2024.03 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLabV2Supervision Type=Fully-supervised, Decoder=C, Backbone=ViT-B2026.05 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RefineNet-LW-152Input Resolution=512x512, FLOPs=71B2018.10 | 82.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BilinearTime (ms)=2.8, Backbone=DINOv2-S/142026.01 | 81.62 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.44 | |
| WeCLIP-FullSupervision Type=Fully-supervised, Decoder=T, Backbone=ViT-B*2026.05 | 81.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoSABackbone=Swin-B, Supervision=I2024.03 | 81.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| †WeCLIP-FullSupervision Type=Fully-supervised, Decoder=C, Backbone=ViT-B*2026.05 | 81.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LiFTParams (M)=1.2, Time (ms)=51.9, Backbone=DINOv2-S/142026.01 | 80.97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.37 | |
| ClearCLIPPublication=ECCV'24, Setting=TF-OVSS2024.08 | 80.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| WideResNetBackbone=ResNet38, Supervision=Fully-supervised2024.02 | 80.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLabv3+Backbone=ResNet-101, Supervision=M2024.03 | 80.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RefineNet-LW-101Input Resolution=512x512, FLOPs=52B2018.10 | 80.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EdgeNeXt-SParams=6.5M, MAdds=8.7G, Framework=DeepLabv3, Input resolution=512x5122022.06 | 80.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GroupViTPublication=CVPR'22, Setting=TL-OVSS2024.08 | 79.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SEREBackbone=ViT-S/162022.06 | 79.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 88.8 | — | — | |
| DHRBackbone=ResNet-101, Supervision=I2024.03 | 79.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MobileViT-SParams=5.7M, MAdds=13.7G, Framework=DeepLabv3, Input resolution=512x5122022.06 | 79.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Weakly-supervised Panoptic SegmentationBackbone=ResNet-101, With COCO annotations=true2018.08 | 79 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 75.7 | 95.8 | — | — | — | — | — | |
| DiCLIPSupervision Type=Weakly-supervised, Decoder=T (Trans. Head), Backbone=ViT-B*2026.05 | 78.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.6 | — | |
| SegformerBackbone=MiT-B1, Supervision=Fully-supervised2024.02 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegformerSupervision Type=Fully-supervised, Decoder=T, Backbone=MiT-B2026.05 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DiCLIPSupervision Type=Weakly-supervised, Decoder=C (Conv. Head), Backbone=ViT-B*2026.05 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.7 | — | |
| RefineNet-LW-50Input Resolution=512x512, FLOPs=33B2018.10 | 78.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NearestTime (ms)=0.6, Backbone=DINOv2-S/142026.01 | 78.29 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 94.37 | |
| ReCoBackbone=DeepLabv3+, Labelled data ratio=Full2021.04 | 77.75 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLab-v2-ResNet-101-CRFInput Resolution=512x5122018.10 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SDIBackbone=ResNet-101, With COCO annotations=true2018.08 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 74.2 | 95.5 | — | — | — | — | — | |
| DeepLabV2Backbone=ResNet101, Supervision=Fully-supervised2024.02 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MARSBackbone=ResNet-101, Supervision=I2024.03 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLabV2Supervision Type=Fully-supervised, Decoder=C, Backbone=RN1012026.05 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLabV3-Res101 (Teacher)Params (M)=61.1M, FLOPs (G)=1294.6G, Crop size=512 x 512, Network Role=Teacher2022.04 | 77.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TCLPublication=CVPR'23, Setting=TL-OVSS2024.08 | 77.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RefineNet-LW-NASNet-MobileInput Resolution=512x512, FLOPs=11.4B2018.10 | 77.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Weakly-supervised Panoptic SegmentationBackbone=ResNet-101, With COCO annotations=false2018.08 | 77.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 74.3 | 96.1 | — | — | — | — | — | |
| FMA-WSSSBackbone=ResNet-101, Supervision=I+C+S2024.03 | 77.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=ViT-S/162022.06 | 77.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 87.5 | — | — | |
| SCLIPPublication=ECCV'24, Setting=TF-OVSS2024.08 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lai et al.Architecture=Deeplabv3+ with ResNet-50 backbone, Labeled-unlabeled ratio=FS, Pre-training=ImageNet2021.04 | 76.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoSABackbone=ResNet-101, Supervision=I2024.03 | 76.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| WeCLIPType=S, Time (min)=270, GPU (G)=6.22026.04 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| WeCLIPSupervision Type=Weakly-supervised, Decoder=T, Backbone=ViT-B*2026.05 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 93.6 | — | |
| MoReSupervision Type=Weakly-supervised, Decoder=C, Backbone=ViT-B2026.05 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.8 | — | |
| Error-corrArchitecture=Deeplabv3+ with ResNet-50 backbone, Labeled-unlabeled ratio=FS, Pre-training=ImageNet2021.04 | 76.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RefineNet-LW-MobileNet-v2Input Resolution=512x512, FLOPs=9.3B2018.10 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Web-DINOArch=7B/14, Probe Type=Linear Probing, Backbone Status=Frozen, Resolution=560x5602026.02 | 76.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| POTType=M, Time (min)=6782026.04 | 76.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SemiSeg-ContrastiveArchitecture=Deeplabv3+ with ResNet-50 backbone, Labeled-unlabeled ratio=FS, Pre-training=ImageNet2021.04 | 75.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| JML-KDBackbone=DL3-R182023.02 | 75.89 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReCLIPPublication=CVPR'24, Setting=C-USS2024.08 | 75.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MobileNet-v2-DeepLab-v3Input Resolution=512x512, FLOPs=5.8B2018.10 | 75.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MobileNetV2Params=4.5M, MAdds=5.8G, Framework=DeepLabv3, Input resolution=512x5122022.06 | 75.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| S4GANBackbone=DeepLabv2, Labelled data ratio=Full2021.04 | 75.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PCRESupervision Type=Weakly-supervised, Decoder=C, Backbone=ViT-B2026.05 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.7 | — | |
| MobileNet-v1-DeepLab-v3Input Resolution=512x512, FLOPs=14.2B2018.10 | 75.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |