Semantic Segmentation on Pascal VOC 21 classes (val)
84.6mIoUODISE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ODISEtraining dataset=COCO Panoptic2023.08 | 84.6 | — | |
| ODISE2023.07 | 84.6 | — | |
| HIPIEBackbone=ViT-H2023.07 | 83.3 | — | |
| ODISEtraining dataset=COCO Panoptic + COCO Caption, variant=caption2023.08 | 82.7 | — | |
| FC-CLIPtraining dataset=COCO Panoptic2023.08 | 81.8 | — | |
| Vision TransformerArchitecture=ViT backbone and linear decoder, Batchsize=16, Supervision=Full supervision (NLL loss)2025.07 | 81.4 | — | |
| HCCE + PCDArchitecture=ViT-linear, Batchsize=16, SL (soft)=true, N=NN, Pretrained=false, Supervision=Scribble2025.07 | 80.94 | — | |
| HCCE + PCDArchitecture=ViT-linear, Batchsize=12, SL (soft)=true, N=NN, Pretrained=false, Supervision=Scribble2025.07 | 80.8 | — | |
| FARCLUSSBackbone=ResNet-101, Ratio=1/2 (732), Params=59.5M2025.06 | 80.3 | — | |
| ESC-NetSize=ViT-B/16, Training Approach=Training-based2024.11 | 80.1 | — | |
| FARCLUSSBackbone=ResNet-101, Ratio=1/4 (366), Params=59.5M2025.06 | 79 | — | |
| FARCLUSSBackbone=ResNet-50, Ratio=Full (1464), Params=59.5M2025.06 | 78.9 | — | |
| DeepLab*Architecture=V3+ (ResNet101), Batchsize=16, Supervision=Full supervision (NLL loss)2025.07 | 78.9 | — | |
| UniMatchBackbone=ResNet-50, Ratio=Full (1464), Venue=CVPR 232025.06 | 78.7 | — | |
| AGMMArchitecture=ViT-linear (Δ), Batchsize=16, SL (soft)=true, Supervision=Scribble2025.07 | 78.7 | — | |
| FARCLUSSBackbone=ResNet-101, Ratio=1/8 (183), Params=59.5M2025.06 | 78.2 | — | |
| HCCE + PCDArchitecture=V3+, Batchsize=16, SL (soft)=true, N=NN, Supervision=Scribble2025.07 | 78.1 | — | |
| HCCE + PCDArchitecture=V3+, Batchsize=12, SL (soft)=true, N=NN, Supervision=Scribble2025.07 | 77.7 | — | |
| FARCLUSSBackbone=ResNet-50, Ratio=1/2 (732), Params=59.5M2025.06 | 77.69 | — | |
| HCCE + PCDArchitecture=V3+, Batchsize=16, SL (soft)=true, N=NN, Pretrained=false, Supervision=Scribble2025.07 | 77.6 | — | |
| HCCE + PQArchitecture=V3+, Batchsize=12, SL (soft)=true, N=NN, Supervision=Scribble2025.07 | 77.5 | — | |
| CAT-SegSize=ViT-B/16, Training Approach=Training-based2024.11 | 77.3 | — | |
| TELArchitecture=V3+, Batchsize=16, SL (soft)=true, Supervision=Scribble2025.07 | 77.1 | — | |
| UniMatchBackbone=ResNet-50, Ratio=1/2 (732), Venue=CVPR 232025.06 | 77.09 | — | |
| HCCE + PCDArchitecture=V3+, Batchsize=12, SL (soft)=true, N=NN, Pretrained=false, Supervision=Scribble2025.07 | 76.7 | — | |
| CorrCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 76.7 | — | |
| DeepLab*Architecture=V3+ (ResNet101), Batchsize=12, Supervision=Full supervision (NLL loss)2025.07 | 76.6 | — | |
| FARCLUSSBackbone=ResNet-50, Ratio=1/4 (366), Params=59.5M2025.06 | 76.5 | — | |
| FARCLUSSBackbone=ResNet-101, Ratio=1/16 (92), Params=59.5M2025.06 | 76.4 | — | |
| CorrMatchBackbone=ResNet-101, Ratio=1/16 (92), Venue=CVPR 242025.06 | 76.4 | — | |
| CorrCLIPSize=ViT-H/14, Training Approach=Training-free2024.11 | 76.4 | — | |
| SEMINARArchitecture=V3+ (Δ), Batchsize=12, GD=true, SL (soft)=true, Supervision=Scribble2025.07 | 76.2 | — | |
| FARCLUSSBackbone=ResNet-50, Ratio=1/8 (183), Params=59.5M2025.06 | 76.18 | — | |
| UniMatchBackbone=ResNet-50, Ratio=1/4 (366), Venue=CVPR 232025.06 | 75.96 | — | |
| DenseCRF loss*Architecture=V3+, Batchsize=12, GD=true, N=DN, Supervision=Scribble2025.07 | 75.8 | — | |
| Diverse CoTBackbone=ResNet-101, Ratio=1/16 (92), Venue=ICCV'232025.06 | 75.7 | — | |
| NonlocalCRF loss*Architecture=V3+, Batchsize=12, GD=true, N=SN, Supervision=Scribble2025.07 | 75.7 | — | |
| HIPIEBackbone=RN502023.07 | 75.7 | — | |
| DeepLabArchitecture=V2 (ResNet101), Batchsize=12, Supervision=Full supervision (NLL loss)2025.07 | 75.6 | — | |
| GridCRF loss*Architecture=V3+, Batchsize=12, N=NN, Supervision=Scribble2025.07 | 75.6 | — | |
| UniMatchBackbone=ResNet-101, Ratio=1/16 (92), Venue=CVPR 232025.06 | 75.2 | — | |
| PSIArchitecture=V3+ (Δ), Batchsize=14, SL (soft)=true, Supervision=Scribble2025.07 | 74.9 | — | |
| CorrCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 74.8 | — | |
| URSSArchitecture=V2 (Δ), Batchsize=16, GD=true, Supervision=Scribble2025.07 | 74.6 | — | |
| SimBaselinetraining dataset=COCO Stuff [5]2023.08 | 74.5 | — | |
| SPMLArchitecture=V2 (Δ), Batchsize=16, GD=true, Supervision=Scribble2025.07 | 74.2 | — | |
| ZegFormertraining dataset=COCO Stuff [5]2023.08 | 73.3 | — | |
| BPGArchitecture=V2 (Δ), Batchsize=10, GD=true, Supervision=Scribble2025.07 | 73.2 | — | |
| FARCLUSSBackbone=ResNet-50, Ratio=1/16 (92), Params=59.5M2025.06 | 72.9 | — | |
| CW-BASSBackbone=ResNet-50, Ratio=1/16 (92), Venue=IJCNN'252025.06 | 72.8 | — | |
| ST++Backbone=ResNet-50, Ratio=1/16 (92), Venue=CVPR 222025.06 | 72.6 | — | |
| UniMatchBackbone=ResNet-50, Ratio=1/8 (183), Venue=CVPR 232025.06 | 72.48 | — | |
| UniMatchBackbone=ResNet-50, Ratio=1/16 (92), Venue=CVPR 232025.06 | 71.9 | — | |
| TridentSize=ViT-H/14, Training Approach=Training-free2024.11 | 70.8 | — | |
| CLIPerSize=ViT-L/14, Training Approach=Training-free2024.11 | 69.8 | — | |
| Supervised OnlyBackbone=ResNet-101, Ratio=1/2 (732)2025.06 | 69.7 | — | |
| CaRSize=ViT-L/14, Training Approach=Training-free2024.11 | 67.6 | — | |
| TridentSize=ViT-B/16, Training Approach=Training-free2024.11 | 67.1 | — | |
| Supervised OnlyBackbone=ResNet-50, Ratio=1/2 (732)2025.06 | 66.7 | — | |
| CLIPerSize=ViT-B/16, Training Approach=Training-free2024.11 | 65.9 | — | |
| CASSSize=ViT-B/16, Training Approach=Training-free2024.11 | 65.8 | — | |
| SC-CLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 65 | — | |
| ProxyCLIPSize=ViT-H/14, Training Approach=Training-free2024.11 | 65 | — | |
| Supervised OnlyBackbone=ResNet-101, Ratio=1/4 (366)2025.06 | 64.8 | — | |
| SC-CLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 64.6 | — | |
| NACLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 64.1 | — | |
| ScribbleSupArchitecture=V2 (VGG16), Batchsize=8, SL (hard)=true, N=DN, Supervision=Scribble2025.07 | 63.1 | — | |
| CLIP-DINOiserSize=ViT-B/16, Training Approach=Training-based2024.11 | 62.1 | — | |
| LaVGSize=ViT-B/16, Training Approach=Training-free2024.11 | 62.1 | — | |
| Supervised OnlyBackbone=ResNet-50, Ratio=1/4 (366)2025.06 | 61.7 | — | |
| ProxyCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 61.3 | — | |
| ResCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 61.1 | — | |
| ProxyCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 60.6 | — | |
| SCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 59.1 | — | |
| CoDeSize=ViT-B/16, Training Approach=Training-based2024.11 | 57.7 | — | |
| FreeDASize=ViT-L/14, Training Approach=Training-free2024.11 | 55.4 | — | |
| Supervised OnlyBackbone=ResNet-101, Ratio=1/8 (183)2025.06 | 55.3 | — | |
| TCLSize=ViT-B/16, Training Approach=Training-based2024.11 | 55 | — | |
| ResCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 54.1 | — | |
| CLIPtraseSize=ViT-B/16, Training Approach=Training-free2024.11 | 53 | — | |
| Supervised OnlyBackbone=ResNet-50, Ratio=1/8 (183)2025.06 | 52.3 | — | |
| ClearCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 51.8 | — | |
| GroupViTtraining dataset=GCC [75]+YFCC [79]2023.08 | 50.7 | — | |
| GroupViT2023.07 | 50.7 | — | |
| LSegtraining dataset=Pascal VOC [27]2023.08 | 47.4 | — | |
| Supervised OnlyBackbone=ResNet-101, Ratio=1/16 (92)2025.06 | 45.1 | — | |
| Supervised OnlyBackbone=ResNet-50, Ratio=1/16 (92)2025.06 | 44 | — | |
| MaskCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 38.8 | — | |
| ZS3Net2023.07 | 38.3 | — | |
| SPNettraining dataset=Pascal VOC [27]2023.08 | 24.3 | — | |
| ZS3Nettraining dataset=Pascal VOC [27]2023.08 | 19.4 | — | |
| CLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 11.5 | — | |
| JAFARBackbone=DINOv2-ViT-S/14, Method Category=Task-Agnostic, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.8444 | 0.9628 | |
| DySampleBackbone=DINOv2-ViT-S/14, Method Category=Task-Dependent, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.8162 | 0.9548 | |
| FeatUpBackbone=DINOv2-ViT-S/14, Method Category=Task-Agnostic, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.8108 | 0.9532 | |
| BilinearBackbone=DINOv2-ViT-S/14, Method Category=Training-free, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.807 | 0.9517 | |
| ReSFuBackbone=DINOv2-ViT-S/14, Method Category=Task-Dependent, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.803 | 0.9505 | |
| CARAFEBackbone=DINOv2-ViT-S/14, Method Category=Task-Dependent, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.8026 | 0.9514 | |
| LIFTBackbone=DINOv2-ViT-S/14, Method Category=Task-Agnostic, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.7806 | 0.9462 | |
| SAPABackbone=DINOv2-ViT-S/14, Method Category=Task-Dependent, Resolution=448x448, Evaluation Protocol=Linear Probing2025.06 | 0.7702 | 0.9407 |