Semantic Segmentation on ADE20K (PixAcc and mIoU)
53.1mIoUDINOv3-L
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DINOv3-LResolution=2242026.05 | 53.1 | 79.2 | |
| ViT-5-Largesegmentation_head=UperNet, parameters=354M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 52 | — | |
| SiameseIM2022.06 | 51.1 | — | |
| RADIOv2.5-LParams (M)=320, Evaluation Protocol=Linear probe2025.08 | 50.68 | — | |
| UniUGG Enc.Params (M)=320, Evaluation Protocol=Linear probe2025.08 | 50.12 | — | |
| AM-RADIO-v2.5Encoder Arch.=ViT-Base, Training Data=DataComp-1B, Training Res.=5122025.03 | 50 | — | |
| DeiT-III-Largesegmentation_head=UperNet, parameters=354M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 49.3 | — | |
| ViT-5-Basesegmentation_head=UperNet, parameters=128M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 49.1 | — | |
| DINOv2-g/14-regParams (M)=1,137, Evaluation Protocol=Linear probe2025.08 | 48.68 | — | |
| MAE2022.06 | 48.1 | — | |
| DeiT-III-Basesegmentation_head=UperNet, parameters=128M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 48 | — | |
| DINO-v2Encoder Arch.=ViT-Large, Training Data=LVD-142M, Training Res.=5182025.03 | 47.7 | — | |
| ViT-5-Smallsegmentation_head=UperNet, parameters=42M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 47.5 | — | |
| MoCo-v32022.06 | 47.3 | — | |
| DINO-v2Encoder Arch.=ViT-Base, Training Data=LVD-142M, Training Res.=5182025.03 | 47.3 | — | |
| MUSEResolution=2562026.05 | 46.5 | 72.8 | |
| CTNetBackbone=ResNet-1012021.04 | 45.94 | — | |
| DUNEEncoder Arch.=ViT-Base, Training Data=DUNE-20.7M, Training Res.=4482025.03 | 45.6 | — | |
| DUNE-B/14-448Params (M)=420, Evaluation Protocol=Linear probe2025.08 | 45.6 | — | |
| CCNetBackbone=ResNet-1012021.04 | 45.22 | — | |
| DeiT-III-Smallsegmentation_head=UperNet, parameters=42M, resolution=512x512, iterations=160k, pre_training=ImageNet-1k2026.02 | 45.2 | — | |
| DUNEEncoder Arch.=ViT-Base, Training Data=DUNE-20.7M, Training Res.=3362025.03 | 44.9 | — | |
| CFNetBackbone=ResNet-1012021.04 | 44.89 | — | |
| PSPNetBackbone=ResNet-1012021.04 | 43.51 | — | |
| AnyUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 42.43 | 75.85 | |
| FeatUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 42.19 | 75.57 | |
| JAFARLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 42.06 | 75.48 | |
| LoftUpLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 42.02 | 75.72 | |
| BilinearLinear probing=1x1 convolution layer, Input resolution=448x448, Upsampling factor=14x or 16x2025.10 | 40.54 | 74.12 | |
| SigLIP-L/14Params (M)=428, Evaluation Protocol=Linear probe2025.08 | 40.53 | — | |
| InternViT-300MResolution=2242026.05 | 40.2 | 61.4 | |
| VTP-L-d64Resolution=2562026.05 | 36.8 | 58.9 | |
| OpenAI CLIP-L/14Params (M)=305, Evaluation Protocol=Linear probe2025.08 | 36.51 | — | |
| MASt3R Enc.Params (M)=303, Evaluation Protocol=Linear probe2025.08 | 32.54 | — | |
| DUSt3R Enc.Params (M)=303, Evaluation Protocol=Linear probe2025.08 | 32.1 | — | |
| SAM-H/16Params (M)=637, Evaluation Protocol=Linear probe2025.08 | 28.08 | — | |
| VA-VAE-d32Resolution=2562026.05 | 19.6 | 43.1 | |
| UniTokResolution=2562026.05 | 19.5 | 43.1 | |
| TokenFlowResolution=2562026.05 | 17.4 | 38.5 | |
| UniLIPResolution=2562026.05 | 15.4 | 35.8 | |
| SD-VAEResolution=2562026.05 | 15.2 | 35.5 | |
| FCN (D-efficient + cosine backbone)Refinement=+ cosine, Backbone Top-1 Acc=77.91%2018.12 | 0.3933 | 0.7925 | |
| FCN (D-efficient + distill w/o mixup backbone)Refinement=+ distill w/o mixup, Backbone Top-1 Acc=78.67%2018.12 | 0.389 | 0.7897 | |
| FCN (D-efficient backbone)Refinement=D-efficient, Backbone Top-1 Acc=77.16%2018.12 | 0.3888 | 0.7888 | |
| FCN (D-efficient + smooth backbone)Refinement=+ smooth, Backbone Top-1 Acc=78.34%2018.12 | 0.3875 | 0.7864 | |
| FCN (D-efficient + mixup w/ distill backbone)Refinement=+ mixup w/ distill, Backbone Top-1 Acc=79.29%2018.12 | 0.384 | 0.7872 | |
| FCN (D-efficient + mixup w/o distill backbone)Refinement=+ mixup w/o distill, Backbone Top-1 Acc=79.16%2018.12 | 0.3799 | 0.7847 | |
| FCN (B-standard backbone)Refinement=B-standard, Backbone Top-1 Acc=76.14%2018.12 | 0.3705 | 0.7808 | |
| CASSBackbone=CLIP ViT-B/162024.11 | — | 48.6 | |
| CLIPtraseBackbone=CLIP ViT-B/162024.11 | — | 38.6 | |
| LaVGBackbone=CLIP ViT-B/162024.11 | — | 37 | |
| NACLIPBackbone=CLIP ViT-B/162024.11 | — | 45.2 | |
| ProxyCLIPBackbone=CLIP ViT-B/162024.11 | — | 49.1 | |
| SCLIPBackbone=CLIP ViT-B/162024.11 | — | 38.7 |