Semantic Segmentation on COCO (mIoU)
67.4mIoUFull-precision
Evaluation Results
| Method | Links | |
|---|---|---|
| Full-precisionBits (W/A)=16/162026.06 | 67.4 | |
| Shift-and-Sum Quantization + LiteVARBits (W/A)=6/62026.06 | 67 | |
| LiteVARBits (W/A)=6/62026.06 | 66.9 | |
| HIPIEBackbone=ViT-H2023.07 | 66.8 | |
| HIPIEBackbone=ViT-H2023.07 | 66.8 | |
| Shift-and-Sum QuantizationBits (W/A)=6/62026.06 | 66.8 | |
| BRECQBits (W/A)=6/62026.06 | 66.5 | |
| Labeled OnlyEncoder=DINOV2-B, #Params=97.5M2024.10 | 66.4 | |
| X-DecoderBackbone=DaViT-B2023.07 | 66 | |
| X-DecoderBackbone=DaViT-B2023.07 | 66 | |
| SEEMBackbone=DaViT-B2023.07 | 65.3 | |
| SEEMBackbone=DaViT-B2023.07 | 65.3 | |
| ODISEBackbone=ViT-H+SD2023.07 | 65.2 | |
| ODISEBackbone=ViT-H+SD2023.07 | 65.2 | |
| Shift-and-Sum Quantization + LiteVARBits (W/A)=4/42026.06 | 65 | |
| Shift-and-Sum QuantizationBits (W/A)=4/42026.06 | 64.6 | |
| LiteVARBits (W/A)=4/42026.06 | 64.4 | |
| BRECQBits (W/A)=4/42026.06 | 63.7 | |
| DaTaSegBackbone=ViTDet-L, Training data=COCO panoptic + ADE semantic + O365 bbox2023.06 | 62.9 | |
| DiveUpV.A. (VFM-agnostic)=true, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 62.71 | |
| DaTaSegBackbone=ViTDet-B, Training data=COCO panoptic + ADE semantic2023.06 | 62.7 | |
| Labeled OnlyEncoder=DINOV2-S, #Params=24.8M2024.10 | 62.5 | |
| X-DecoderBackbone=FocalT2023.07 | 62.4 | |
| X-DecoderBackbone=FocalT2023.07 | 62.4 | |
| LoftUpV.A. (VFM-agnostic)=false, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 62.23 | |
| NAFV.A. (VFM-agnostic)=true, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=true2026.03 | 62.18 | |
| SEEMBackbone=FocalT2023.07 | 61.2 | |
| SEEMBackbone=FocalT2023.07 | 61.2 | |
| AnyUp-SV.A. (VFM-agnostic)=true, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 61.1 | |
| JAFARV.A. (VFM-agnostic)=false, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 60.78 | |
| FeatUpV.A. (VFM-agnostic)=false, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 60.1 | |
| HIPIEBackbone=RN502023.07 | 59.5 | |
| HIPIEBackbone=RN502023.07 | 59.5 | |
| BilinearV.A. (VFM-agnostic)=true, Backbone=DINOv2-S, Evaluation Protocol=Linear Probing, Refinement stage=false2026.03 | 59.03 | |
| DaTaSegBackbone=R-50, Training data=COCO panoptic2023.06 | 57.7 | |
| ODISEBackbone=UNet+M2F, Training data=LAION+CLIP+COCO2023.06 | 52.4 | |
| UniMatch + ECOCSegLabel Ratio=1/322025.12 | 51.6 | |
| AllSparkLabel Ratio=1/322025.12 | 50.9 | |
| DiGSeginput size=512²2026.04 | 50.8 | |
| UniMatch + ECOCSegLabel Ratio=1/642025.12 | 49.9 | |
| UniMatchLabel Ratio=1/322025.12 | 49.8 | |
| AllSparkLabel Ratio=1/642025.12 | 49.5 | |
| EoMTinput size=512²2026.04 | 48.7 | |
| MSegBackbone=HRNet-W48, Venue=CVPR 20, Label space=MR2024.07 | 48.6 | |
| UniMatchLabel Ratio=1/642025.12 | 48.2 | |
| SegMANinput size=512²2026.04 | 48.2 | |
| 4M-21Model Scale=XL, Evaluation Protocol=Out-of-the-box2024.06 | 48.1 | |
| VWFormer-B5input size=512²2026.04 | 48 | |
| AutoUniSeg (Ours)Backbone=HRNet-W48, Label space=Auto2024.07 | 46.7 | |
| SegFormer-B5input size=512²2026.04 | 46.7 | |
| 4M-21Model Scale=Large, Evaluation Protocol=Out-of-the-box2024.06 | 46.4 | |
| UniMatch + ECOCSegLabel Ratio=1/1282025.12 | 46.2 | |
| PC2SegLabel Ratio=1/322025.12 | 46.1 | |
| CGRSeg-Linput size=512²2026.04 | 46 | |
| OffSeg-Linput size=512²2026.04 | 46 | |
| AllSparkLabel Ratio=1/1282025.12 | 45.4 | |
| UniMatchLabel Ratio=1/1282025.12 | 44.4 | |
| Unified-IOModel Scale=XL, Evaluation Protocol=Out-of-the-box2024.06 | 44.3 | |
| PC2SegLabel Ratio=1/642025.12 | 43.7 | |
| PseudoSegLabel Ratio=1/322025.12 | 43.6 | |
| 4M-21Model Scale=Base, Evaluation Protocol=Out-of-the-box2024.06 | 42.5 | |
| Sup-onlyLabel Ratio=1/322025.12 | 42.2 | |
| PseudoSegLabel Ratio=1/642025.12 | 41.8 | |
| UniMatch + ECOCSegLabel Ratio=1/2562025.12 | 41.8 | |
| Unified-IO 2Model Scale=XXL, Evaluation Protocol=Out-of-the-box2024.06 | 41.7 | |
| AllSparkLabel Ratio=1/2562025.12 | 41.6 | |
| Unified-IOModel Scale=Large, Evaluation Protocol=Out-of-the-box2024.06 | 41.6 | |
| PC2SegLabel Ratio=1/1282025.12 | 40.1 | |
| Unified-IO 2Model Scale=XL, Evaluation Protocol=Out-of-the-box2024.06 | 39.7 | |
| Uni NLL+Backbone=SNp-DN161, Venue=IJCV 24, Label space=MC2024.07 | 39.3 | |
| PseudoSegLabel Ratio=1/1282025.12 | 39.1 | |
| UniMatchLabel Ratio=1/2562025.12 | 38.9 | |
| Unified-IO 2Model Scale=Large, Evaluation Protocol=Out-of-the-box2024.06 | 38.9 | |
| OpenSegBackbone=Eff-b7, Training data=COCO+Loc. Narr.2023.06 | 38 | |
| Single datasetBackbone=HRNet-W48, Label space=DS2024.07 | 38 | |
| Sup-onlyLabel Ratio=1/642025.12 | 37.8 | |
| PC2SegLabel Ratio=1/2562025.12 | 37.5 | |
| PseudoSegLabel Ratio=1/2562025.12 | 37.1 | |
| Multi-SegHeadBackbone=HRNet-W48, Label space=DS2024.07 | 36.7 | |
| OpenSegBackbone=R-101, Training data=COCO pan+cap2023.06 | 36.1 | |
| Auto univ.Backbone=SNp-RN18, Venue=BMVC 22, Label space=Auto2024.07 | 35.6 | |
| NLL+Backbone=SNp-RN18, Venue=WACV 22, Label space=MC2024.07 | 35.4 | |
| CLS-SEGMethod Category=CLIP-based, Backbone=ViT-B/16, Post-processing=denseCRF2023.12 | 35.3 | |
| UniMatch + ECOCSegLabel Ratio=1/5122025.12 | 34.5 | |
| AllSparkLabel Ratio=1/5122025.12 | 34.1 | |
| CLS-SEGMethod Category=CLIP-based, Backbone=ViT-B/162023.12 | 34 | |
| Sup-onlyLabel Ratio=1/1282025.12 | 33.6 | |
| Unified-IOModel Scale=Base, Evaluation Protocol=Out-of-the-box2024.06 | 32.9 | |
| UniMatchLabel Ratio=1/5122025.12 | 31.9 | |
| PC2SegLabel Ratio=1/5122025.12 | 29.9 | |
| PseudoSegLabel Ratio=1/5122025.12 | 29.8 | |
| OpenSDTraining Data=ADE, Backbone=Swin-T2023.12 | 29.1 | |
| Sup-onlyLabel Ratio=1/2562025.12 | 28 | |
| NamedMaskMethod Category=CLIP-based2023.12 | 27.7 | |
| SegCLIPArch.=ViT, Init.=true, Training Data=CC+COCO, Sup.=Text, Zero-Shot=true2022.11 | 26.5 | |
| SegCLIPMethod Category=CLIP-based2023.12 | 26.5 | |
| CLIPpyArch=ViT, Dataset=HQITP-134M, SSP=true2022.10 | 25.5 | |
| CLIPSurgeryMethod Category=CLIP-based, Implementation=Re-implemented2023.12 | 25.2 | |
| GroupViTArch.=ViT, Init.=false, Training Data=CC12M+YFCC, Sup.=Text, Zero-Shot=true2022.11 | 24.3 | |
| GroupViTMethod Category=CLIP-based2023.12 | 24.3 |