Semantic Segmentation on COCO Stuff
3,100mIoUCLS-SEG
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CLS-SEGMethod Category=CLIP-based, Backbone=ViT-B/16, Post-processing=denseCRF2023.12 | 3,100 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLS-SEGMethod Category=CLIP-based, Backbone=ViT-B/162023.12 | 3,010 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPSurgeryMethod Category=CLIP-based, Implementation=Re-implemented2023.12 | 2,970 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReCoMethod Category=CLIP-based2023.12 | 2,630 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskCLIPMethod Category=CLIP-based, Implementation=Re-implemented2023.12 | 2,390 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransFGUMethod Category=Vanilla USS2023.12 | 1,750 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PiCIE+HMethod Category=Vanilla USS2023.12 | 1,440 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PiCIEMethod Category=Vanilla USS2023.12 | 1,380 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IICMethod Category=Vanilla USS2023.12 | 670 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RaysUp2026.06 | 62.32 | — | — | — | — | — | — | — | — | — | 81.47 | — | — | |
| LoftUp2026.06 | 62.23 | — | — | — | — | — | — | — | — | — | 81.38 | — | — | |
| AnyUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 62.16 | — | — | — | — | — | — | — | — | — | 81.37 | — | — | |
| LoftUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 62.15 | — | — | — | — | — | — | — | — | — | 81.32 | — | — | |
| AnyUp2026.06 | 62.14 | — | — | — | — | — | — | — | — | — | 81.38 | — | — | |
| FeatUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 61.95 | — | — | — | — | — | — | — | — | — | 81.14 | — | — | |
| FeatUp2026.06 | 61.89 | — | — | — | — | — | — | — | — | — | 81.1 | — | — | |
| JAFARLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 61.82 | — | — | — | — | — | — | — | — | — | 81.07 | — | — | |
| JAFAR2026.06 | 61.79 | — | — | — | — | — | — | — | — | — | 81.11 | — | — | |
| Bilinear2026.06 | 59.58 | — | — | — | — | — | — | — | — | — | 79.42 | — | — | |
| BilinearLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 59.48 | — | — | — | — | — | — | — | — | — | 79.32 | — | — | |
| DeiT-SBackbone=DeiT-S, #Param. (M)=52M2026.03 | 54.23 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskFormer Swin LBackbone=Swin L2026.04 | 52.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Youtu-VL (4B)Additions=None2026.01 | 52.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lorentz (Mask2Former SL)Backbone=Swin L2026.04 | 51.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskFormer Swin BBackbone=Swin B2026.04 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lorentz (MaskFormer SL)Backbone=Swin L2026.04 | 50.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiT-HModel Size=Huge, Training Protocol=universal2024.03 | 49.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiTAdditions=Parallel Decoding2026.01 | 49.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lorentz (MaskFormer SB)Backbone=Swin B2026.04 | 48.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv3-B/16Protocol=Linear probing, Backbone=B/162026.05 | 48.41 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=642026.05 | 47.92 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LawinBackbone=MiT-B5, FLOPs(G)=942022.01 | 47.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lorentz (SegFormer-B3)Backbone=B32026.04 | 47.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AM-RADIOv2.5-B/16Protocol=Linear probing, Backbone=B/162026.05 | 46.96 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegFormerBackbone=MiT-B5, FLOPs(G)=1122022.01 | 46.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Segformer (MiT-B5)Additions=MLP Decoder2026.01 | 46.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegFormerSteps=12025.06 | 46.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegFormer (B4)Backbone=B42026.04 | 46.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiT-LModel Size=Large, Training Protocol=universal2024.03 | 46 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SETR-MLABackbone=ViT-L, Pre-trained=ImageNet22K2022.01 | 45.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAN (ViT-L)Additions=Decoupled Head2026.01 | 45.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegFormer (B3)Backbone=B32026.04 | 45.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-reg-B/14Protocol=Linear probing, Backbone=B/142026.05 | 45.42 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OursBackbone=SDv1.4, Resolution=512 x 5122026.06 | 45 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlowSegBackbone=(PixNerd), Pretrain Data=IN-1k2026.03 | 44.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv2-B/14Protocol=Linear probing, Backbone=B/142026.05 | 44.85 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SegFormerBackbone=MiT-B2, Pretrain Data=IN-1k2026.03 | 44.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=82026.05 | 43.85 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DiffSegBackbone=SDv1.4, Resolution=512 x 5122026.06 | 43.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiT-BModel Size=Base, Training Protocol=universal2024.03 | 42.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OCRNetBackbone=HRNetW48, FLOPs(G)=1652022.01 | 42.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OCRNetBackbone=HRNet-W48, Pretrain Data=IN-1k2026.03 | 42.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskFormerTask=SS2024.05 | 41.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskCut2026.06 | 41.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lorentz (DeepLab-R101)Backbone=R1012026.04 | 41.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OCR (Seg. transformer)Baseline=HRNetV2-W48, Stride=4x, Context schemes=R2019.09 | 40.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv3 + Recursive-NCut2026.06 | 40.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ACNetBaseline=ResNet-101 + MG, Stride=4x, Context schemes=M,R2019.09 | 40.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EMANetBaseline=ResNet-101 + MG, Stride=8x, Context schemes=R2019.09 | 39.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DANetBaseline=ResNet-101 + MG, Stride=8x, Context schemes=R2019.09 | 39.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SVCNetBaseline=ResNet-101, Stride=4x, Context schemes=R2019.09 | 39.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SymmFlowBackbone=(SD2.1), Pretrain Data=LSTI2026.03 | 39.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SymmFlowSteps=252025.06 | 39.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OCR (Seg. transformer)Baseline=ResNet-101, Stride=8x, Context schemes=R2019.09 | 39.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SGRBaseline=ResNet-101 + ASPP, Stride=8x, Context schemes=R2019.09 | 39.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| POMPmode=open-vocabulary2023.04 | 39.1 | — | — | — | 39.9 | — | — | 38.2 | — | — | — | — | — | |
| DSSPNBaseline=ResNet-101, Stride=8x2019.09 | 38.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SigLIP 2-B/16Protocol=Linear probing, Backbone=B/162026.05 | 38.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EQ-VMamba-TBackbone=EQ-VMamba-T, #Param. (M)=18M2026.03 | 38.69 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VMamba-TBackbone=VMamba-T, #Param. (M)=62M2026.03 | 38.67 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SemFlowTask=SS, Sampler=Euler-252024.05 | 38.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeeplabV3+Backbone=ResNet101, FLOPs(G)=2552022.01 | 38.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeeplabV3+Backbone=ResNet50, Pretrain Data=IN-1k2026.03 | 38.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EQ-VMamba-SBackbone=EQ-VMamba-S, #Param. (M)=25M2026.03 | 38.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskCLIP+paradigm=transductive2021.12 | 38.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Fully Sup.2021.12 | 38.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepLabV32026.04 | 38 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NonLocalBackbone=ResNet101, FLOPs(G)=2782022.01 | 37.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ZSSegmode=open-vocabulary2023.04 | 37.8 | — | — | — | 39.3 | — | — | 36.3 | — | — | — | — | — | |
| VMamba-SBackbone=VMamba-S, #Param. (M)=82M2026.03 | 37.43 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaskFormerSteps=12025.06 | 37.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MSVMamba-TBackbone=MSVMamba-T, #Param. (M)=65M2026.03 | 36.63 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 36.18 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MSVMamba-MBackbone=MSVMamba-M, #Param. (M)=42M2026.03 | 35.94 | — | — | — | — | — | — | — | — | — | — | — | — | |
| XCiT-M24Backbone=XCiT-M24, #Param. (M)=112M2026.03 | 35.75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SemFlowSteps=252025.06 | 35.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CaGNetparadigm=transductive2021.12 | 35.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ProMerge2026.06 | 35.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CaGNetparadigm=inductive2021.12 | 35.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| STRICTparadigm=transductive2021.12 | 35.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DFNCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 35.23 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SPNet-Cparadigm=inductive2021.12 | 35.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| XCiT-S24Backbone=XCiT-S24, #Param. (M)=76M2026.03 | 35.11 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-SBackbone=ConvNeXt-S, #Param. (M)=70M2026.03 | 35.03 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ZS3Netparadigm=transductive2021.12 | 34.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ZegFormermode=open-vocabulary2023.04 | 34.8 | — | — | — | 36.6 | — | — | 33.2 | — | — | — | — | — | |
| ZS3Netparadigm=inductive2021.12 | 34.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SPNetparadigm=inductive2021.12 | 34.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SPNetparadigm=transductive2021.12 | 34.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 34.47 | — | — | — | — | — | — | — | — | — | — | — | — |