Semantic Segmentation on PASCAL VOC 2012
87.5mIoULingBot-Vision ViT-g
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LingBot-Vision ViT-gNumber of Parameters=1B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 87.5 | — | — | — | — | — | |
| DINOv3Param.=7B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 86.6 | — | — | — | — | — | |
| DINOv3Number of Parameters=7B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 86.6 | — | — | — | — | — | |
| DINOv3 ViT-H+Param.=0.8B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 85.8 | — | — | — | — | — | |
| DINOv3 ViT-H+Number of Parameters=0.8B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 85.8 | — | — | — | — | — | |
| AM-RADIOv2.5Param.=1B, Patch Size=14, Evaluation Protocol=Linear probe, Input Resolution=448x4482026.03 | 85.4 | — | — | — | — | — | |
| AM-RADIOv2.5Number of Parameters=1B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Input Resolution=448 x 4482026.07 | 85.4 | — | — | — | — | — | |
| V-JEPA 2.1 ViT-GParam.=2B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 85 | — | — | — | — | — | |
| V-JEPA 2.1 ViT-GNumber of Parameters=2B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 85 | — | — | — | — | — | |
| V-JEPA 2.1 ViT-gParam.=1B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 84.7 | — | — | — | — | — | |
| V-JEPA 2.1 ViT-gNumber of Parameters=1B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 84.7 | — | — | — | — | — | |
| DecoupleCSSSetting=15-1, Foundation Model=SAM (~632M params)2026.03 | 83.8 | — | — | — | — | — | |
| DINOv2Param.=1B, Patch Size=14, Evaluation Protocol=Linear probe, Input Resolution=448x4482026.03 | 83.1 | — | — | — | — | — | |
| DINOv2Number of Parameters=1B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Input Resolution=448 x 4482026.07 | 83.1 | — | — | — | — | — | |
| DecoupleCSSSetting=19-1, Foundation Model=SAM (~632M params)2026.03 | 82.9 | — | — | — | — | — | |
| PEspatialParam.=2B, Patch Size=14, Evaluation Protocol=Linear probe, Input Resolution=448x4482026.03 | 82.7 | — | — | — | — | — | |
| PEspatialNumber of Parameters=2B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Input Resolution=448 x 4482026.07 | 82.7 | — | — | — | — | — | |
| PS-MTArchitecture=DeeplabV3+, Backbone=ResNet1012021.11 | 81.19 | — | — | — | — | — | |
| PS-MTArchitecture=DeeplabV3+, Backbone=ResNet1012021.11 | 80.01 | — | — | — | — | — | |
| PS-MTArchitecture=DeeplabV3+, Backbone=ResNet502021.11 | 78.71 | — | — | — | — | — | |
| PS-MTLabeled samples=7322021.11 | 78.42 | — | — | — | — | — | |
| PS-MTArchitecture=DeeplabV3+, Backbone=ResNet502021.11 | 78.08 | — | — | — | — | — | |
| AdvCAMArchitecture=Deeplabv2, Backbone=ResNet1012021.11 | 77.8 | — | — | — | — | — | |
| Joint2021.06 | 77.43 | 7,977 | 72.35 | — | — | — | |
| Joint2021.06 | 77.43 | — | — | 77.51 | 77.04 | — | |
| DnC-4.5kPre-training dataset=JFT-300M, Training iterations=4.5k2021.05 | 76.9 | — | — | — | — | — | |
| DnC-4.5kPre-training dataset=YFCC, Training iterations=4.5k2021.05 | 76.6 | — | — | — | — | — | |
| PS-MTLabeled samples=3662021.11 | 76.57 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/22021.06 | 76.5 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/22021.06 | 76.5 | — | — | — | — | — | |
| SSUL-Mexemplar-memory=true2021.06 | 76.49 | — | — | 77.83 | 49.76 | — | |
| BYOLProtocol=Fine-tuning, Backbone=ResNet-502021.06 | 76.3 | — | — | — | — | — | |
| MoCLR-5kPre-training dataset=JFT-300M, Training iterations=5k2021.05 | 76.1 | — | — | — | — | — | |
| CACArchitecture=DeeplabV3+, Backbone=ResNet502021.11 | 76.1 | — | — | — | — | — | |
| Web-DINOParam.=7B, Patch Size=14, Evaluation Protocol=Linear probe, Input Resolution=448x4482026.03 | 76.1 | — | — | — | — | — | |
| Web-DINONumber of Parameters=7B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Input Resolution=448 x 4482026.07 | 76.1 | — | — | — | — | — | |
| SSL-HSIC (w/ target)Protocol=Fine-tuning, Backbone=ResNet-50, Target network=true2021.06 | 76 | — | — | — | — | — | |
| CPSLabeled samples=7322021.11 | 75.88 | — | — | — | — | — | |
| BYOL-5kPre-training dataset=JFT-300M, Training iterations=5k2021.05 | 75.8 | — | — | — | — | — | |
| PS-MTArchitecture=PSPNet, Backbone=ResNet502021.11 | 75.74 | — | — | — | — | — | |
| BYOL-5kPre-training dataset=YFCC, Training iterations=5k2021.05 | 75.5 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/42021.06 | 75.5 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/42021.06 | 75.5 | — | — | — | — | — | |
| CogCaSSetting=15-12026.03 | 75.5 | — | — | — | — | — | |
| SSUL2021.06 | 75.44 | — | — | 77.73 | 29.68 | — | |
| SimCLRProtocol=Fine-tuning, Backbone=ResNet-502021.06 | 75.2 | — | — | — | — | — | |
| MoCLR-5kPre-training dataset=YFCC, Training iterations=5k2021.05 | 75.1 | — | — | — | — | — | |
| Yuan et al.Architecture=DeeplabV3+, Backbone=ResNet1012021.11 | 75 | — | — | — | — | — | |
| SSL-HSIC (w/o target)Protocol=Fine-tuning, Backbone=ResNet-50, Target network=false2021.06 | 74.9 | — | — | — | — | — | |
| PS-MTArchitecture=PSPNet, Backbone=ResNet502021.11 | 74.59 | — | — | — | — | — | |
| CACArchitecture=DeeplabV3+, Backbone=ResNet502021.11 | 74.5 | — | — | — | — | — | |
| ImageNet Super.Pre-training dataset=ImageNet, Supervision=Supervised2021.05 | 74.4 | — | — | — | — | — | |
| Supervised-INProtocol=Fine-tuning, Backbone=ResNet-502021.06 | 74.4 | — | — | — | — | — | |
| KDEPData=10%, Epoch=900, Time (/h)=40, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 74.28 | — | — | — | — | — | |
| CPCLNetwork=ResNet-502022.11 | 74.25 | — | — | — | — | — | |
| GCTLabeled data ratio=1/22021.06 | 74 | — | — | — | — | — | |
| GCTLabeled data ratio=1/22021.06 | 74 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/22021.06 | 73.9 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/22021.06 | 73.9 | — | — | — | — | — | |
| DARSArchitecture=PSPNet, Backbone=ResNet502021.11 | 73.89 | — | — | — | — | — | |
| KDEPData=100%, Epoch=18, Time (/h)=8, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 73.82 | — | — | — | — | — | |
| PseudoSegArchitecture=DeeplabV3+, Backbone=ResNet502021.11 | 73.8 | — | — | — | — | — | |
| KDEPData=100%, Epoch=90, Time (/h)=40, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 73.75 | — | — | — | — | — | |
| PLOP2021.06 | 73.54 | — | — | 75.35 | 37.35 | — | |
| SupervisedAnnotation=Fine, Architecture=SegFormer (MiT-B0)2026.04 | 73.5 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/82021.06 | 73.4 | — | — | — | — | — | |
| GuidedMix-NetLabeled data ratio=1/82021.06 | 73.4 | — | — | — | — | — | |
| AdvSSLLabeled data ratio=1/22021.06 | 73.3 | — | — | — | — | — | |
| AdvSSLLabeled data ratio=1/22021.06 | 73.3 | — | — | — | — | — | |
| MTLabeled data ratio=1/22021.06 | 73.2 | — | — | — | — | — | |
| MTLabeled data ratio=1/22021.06 | 73.2 | — | — | — | — | — | |
| CCTArchitecture=PSPNet, Backbone=ResNet502021.11 | 73.2 | — | — | — | — | — | |
| PseudoSegArchitecture=DeeplabV3+, Backbone=ResNet1012021.11 | 73.2 | — | — | — | — | — | |
| SP. o.Data=100%, Epoch=90, Time (/h)=39, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 73.13 | — | — | — | — | — | |
| KDEPData=10%, Epoch=180, Time (/h)=8, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 72.82 | — | — | — | — | — | |
| GCTLabeled data ratio=1/42021.06 | 72.8 | — | — | — | — | — | |
| GCTLabeled data ratio=1/42021.06 | 72.8 | — | — | — | — | — | |
| SigLIP 2Param.=1B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 72.7 | — | — | — | — | — | |
| SigLIP 2Number of Parameters=1B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 72.7 | — | — | — | — | — | |
| KDEPData=100%, Epoch=9, Time (/h)=4, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 72.43 | — | — | — | — | — | |
| PseudoSegLabeled samples=7322021.11 | 72.41 | — | — | — | — | — | |
| PseudoSegNetwork=ResNet-1012022.11 | 72.41 | — | — | — | — | — | |
| KDEPData=10%, Epoch=90, Time (/h)=4, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 72.34 | — | — | — | — | — | |
| CPCLNetwork=ResNet-502022.11 | 72.14 | — | — | — | — | — | |
| FP BASELINEArchitecture=DeepLabv32024.05 | 72.1 | — | — | — | — | — | |
| S4LLabeled data ratio=1/22021.06 | 72 | — | — | — | — | — | |
| S4LLabeled data ratio=1/22021.06 | 72 | — | — | — | — | — | |
| CPSLabeled samples=3662021.11 | 71.71 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/42021.06 | 71.7 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/42021.06 | 71.7 | — | — | — | — | — | |
| SeSAMAnnotation=Scribble, Architecture=SegFormer (MiT-B0)2026.04 | 71.4 | — | — | — | — | — | |
| SSUL-Mexemplar-memory=true2021.06 | 71.37 | 7,836 | 49.01 | — | — | — | |
| KDEPData=100%, Epoch=90, Time (/h)=432022.03 | 71.07 | — | — | — | — | — | |
| SP. b.Data=100%, Epoch=18, Time (/h)=7.8, Teacher Backbone=ResNet-50, Student Backbone=ResNet-18, Evaluation Protocol=fine-tuned2022.03 | 71.02 | — | — | — | — | — | |
| MTLabeled data ratio=1/42021.06 | 70.9 | — | — | — | — | — | |
| MTLabeled data ratio=1/42021.06 | 70.9 | — | — | — | — | — | |
| AdvSSLLabeled data ratio=1/42021.06 | 70.8 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/82021.06 | 70.8 | — | — | — | — | — | |
| AdvSSLLabeled data ratio=1/42021.06 | 70.8 | — | — | — | — | — | |
| CutMixLabeled data ratio=1/82021.06 | 70.8 | — | — | — | — | — |