Semantic Segmentation on ADE20K (mIoU)
64.01mIoUCM-GLasso
Evaluation Results
| Method | Links | |
|---|---|---|
| CM-GLasso2026.04 | 64.01 | |
| DINOv3 + SpatialBoostevaluation_protocol=multi-scale evaluation2026.03 | 63.1 | |
| InternImage-H2026.04 | 62.9 | |
| DINOv3evaluation_protocol=multi-scale evaluation2026.03 | 60.3 | |
| DINOv3 + SpatialBoostevaluation_protocol=linear probing2026.03 | 59.7 | |
| TC-JEPABackbone=ViT-L/14, Pre-training Data=CC27M, Protocol=supervised finetuning2026.05 | 58.8 | |
| DiGSeginput size=512²2026.04 | 58.6 | |
| OmniVec22026.04 | 58.5 | |
| Mask2Former-Swin-Linput size=640²2026.04 | 57.3 | |
| EoMTinput size=512²2026.04 | 57.1 | |
| OneFormer2026.04 | 57 | |
| OneFormerinput size=640²2026.04 | 57 | |
| MhSAViT=Large, Attention=MhSA2026.05 | 56.84 | |
| TC-JEPABackbone=ViT-B/16, Pre-training Data=CC27M, Protocol=supervised finetuning2026.05 | 56.8 | |
| SoftmaxBackbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=310.60, Peak Mem. (GB)=1.3181, Throughput (imgs/s)=36.52, Segmentation Head=Mask2former2026.03 | 56.73 | |
| Lorentz (Mask2Former SL)Backbone=Swin L2026.04 | 56.7 | |
| Mask2Former Swin LBackbone=Swin L2026.04 | 56.1 | |
| DW[all]ViT=Large, Attention=DW[all], |S|=12/242026.05 | 56.05 | |
| DINOv3Param.=7B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 55.9 | |
| DINOv3evaluation_protocol=linear probing2026.03 | 55.9 | |
| DINOv3Number of Parameters=7B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 55.9 | |
| TC-JEPAViT=H/14, Protocol=FT2026.05 | 55.7 | |
| ViT-AdaLA (Ours)Backbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2163↓7.7%, Throughput (imgs/s)=41.56↑16.1%, Segmentation Head=Mask2former2026.03 | 55.55 | |
| TC-JEPABackbone=ViT-B/16, Pre-training Data=YFCC15M, Protocol=supervised finetuning2026.05 | 55.2 | |
| DINOv3Size=L2026.07 | 55 | |
| DINOv2 + SpatialBoostevaluation_protocol=multi-scale evaluation2026.03 | 54.9 | |
| DINOv3 ViT-H+Param.=0.8B, Patch Size=16, Evaluation Protocol=Linear probe, Input Resolution=512x5122026.03 | 54.8 | |
| DINOv3 ViT-H+Number of Parameters=0.8B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 54.8 | |
| VWFormer-B5input size=512²2026.04 | 54.7 | |
| Lorentz (MaskFormer SL)Backbone=Swin L2026.04 | 54.5 | |
| SoftmaxBackbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=312.46, Peak Mem. (GB)=1.3636, Throughput (imgs/s)=41.05, Segmentation Head=Mask2former2026.03 | 54.4 | |
| data2vecViT=L/16, Protocol=FT2026.05 | 54.4 | |
| TIPSevaluation_protocol=multi-scale evaluation2026.03 | 54.1 | |
| MaskFormer Swin LBackbone=Swin L2026.04 | 54.1 | |
| SPARCBackbone=ViT-B/16, Pre-training Data=CC27M, Protocol=supervised finetuning2026.05 | 54 | |
| SoftmaxBackbone=DINOv2-B2026.03 | 53.93 | |
| MhSAViT=Base, Attention=MhSA2026.05 | 53.83 | |
| V-JEPAv2evaluation_protocol=multi-scale evaluation2026.03 | 53.8 | |
| MaskDistill*Teacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=3002026.06 | 53.8 | |
| LingBot-Vision ViT-gNumber of Parameters=1B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Input Resolution=512 x 5122026.07 | 53.5 | |
| StoPViT=H/14, Protocol=FT2026.05 | 53.3 | |
| SegMANinput size=512²2026.04 | 53.2 | |
| ViT-AdaLA (Ours)Backbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=264.04↓15.5%, Peak Mem. (GB)=1.2620↓7.4%, Throughput (imgs/s)=47.05↑14.6%, Segmentation Head=Mask2former2026.03 | 53.16 | |
| AM-RADIOv2.5Param.=1B, Patch Size=14, Evaluation Protocol=Linear probe, Input Resolution=448x4482026.03 | 53 | |
| DINOv2evaluation_protocol=multi-scale evaluation2026.03 | 53 | |
| AM-RADIOv2.5Number of Parameters=1B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Input Resolution=448 x 4482026.07 | 53 | |
| CAE v2Teacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=3002026.06 | 52.9 | |
| dino.txtevaluation_protocol=multi-scale evaluation2026.03 | 52.8 | |
| AdaMerging+TAPBase Model=DINOv2, Selection mechanism=TAP2026.04 | 52.8 | |
| MacFormerYear=2024, Backbone=MiT-B5, Params. (M)=103.0, GFLOPs=152.42024.11 | 52.8 | |
| ExPLoRe (64-exp)Teacher=CLIP-B, Params=1.86B, GFLOPs=13.86, Epochs=3002026.06 | 52.8 | |
| LingBot-VisionSize=L2026.07 | 52.75 | |
| SCASegBackbone=MiT-B5, Params. (M)=92.4, GFLOPs=88.92024.11 | 52.7 | |
| BEiT v2Teacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=3002026.06 | 52.7 | |
| MILAN†Teacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=4002026.06 | 52.7 | |
| ViT-AdaLA (Stage 2)Backbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2163↓7.7%, Throughput (imgs/s)=41.56↑16.1%, Segmentation Head=Mask2former2026.03 | 52.46 | |
| EUPE-ViT-BBackbone=ViT-B2026.03 | 52.4 | |
| DreamLIPBackbone=ViT-B/16, Pre-training Data=CC30M, Protocol=supervised finetuning2026.05 | 52.4 | |
| MVPTeacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=3002026.06 | 52.4 | |
| MTA-clip (ViT-B)Backbone=ViT-B2026.04 | 52.3 | |
| SPARCBackbone=ViT-B/16, Pre-training Data=YFCC15M, Protocol=supervised finetuning2026.05 | 52.3 | |
| SegmenterModel Scale=Large, GELU Type=GELU, Mode=Baseline2025.09 | 52.29 | |
| LDMSeginput size=512²2026.04 | 52.2 | |
| DINOv2 + SpatialBoostevaluation_protocol=linear probing2026.03 | 52 | |
| VWFormerYear=2024, Backbone=MiT-B5, Params. (M)=84.6, GFLOPs=96.12024.11 | 52 | |
| U-MixFormerYear=2025, Backbone=MiT-B5, Params. (M)=93.0, GFLOPs=149.52024.11 | 51.9 | |
| DINOv3-ViT-BBackbone=ViT-B2026.03 | 51.8 | |
| Lorentz (MaskFormer SB)Backbone=Swin B2026.04 | 51.8 | |
| SegFormer-B5input size=640²2026.04 | 51.8 | |
| DINOv3Size=B2026.07 | 51.74 | |
| MTA-clip (MiT-B4)Backbone=MiT-B42026.04 | 51.7 | |
| DW[all]ViT=Base, Attention=DW[all], |S|=6/122026.05 | 51.7 | |
| TIPSv2Backbone=g/142026.04 | 51.6 | |
| TC-JEPAViT=B/16, Protocol=FT2026.05 | 51.6 | |
| ViT-AdaLABackbone=DINOv2-B, Pretraining Epochs=402026.03 | 51.46 | |
| LingBot-VisionSize=B2026.07 | 51.44 | |
| MetaSegYear=2024, Backbone=MiT-B5, Params. (M)=85.0, GFLOPs=74.52024.11 | 51.4 | |
| ViT-AdaLA (Stage 2)Backbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=264.04↓15.5%, Peak Mem. (GB)=1.2620↓7.4%, Throughput (imgs/s)=47.05↑14.6%, Segmentation Head=Mask2former2026.03 | 51.39 | |
| DINOv3-B/16Protocol=Linear probing, Backbone=B/162026.05 | 51.35 | |
| V-JEPAv2evaluation_protocol=linear probing2026.03 | 51.3 | |
| FeedFormerYear=2023, Backbone=MiT-B5, Params. (M)=85.6, GFLOPs=79.82024.11 | 51.2 | |
| I-JEPAViT=H/14, Protocol=FT2026.05 | 51.2 | |
| MaskFormer Swin BBackbone=Swin B2026.04 | 51.1 | |
| SegFormerYear=2021, Backbone=MiT-B5, Params. (M)=84.7, GFLOPs=183.32024.11 | 51 | |
| MacFormerYear=2024, Backbone=MiT-B4, Params. (M)=82.0, GFLOPs=76.72024.11 | 50.9 | |
| SCASegBackbone=MiT-B4, Params. (M)=71.8, GFLOPs=72.92024.11 | 50.9 | |
| MAEViT=H/14, Protocol=FT2026.05 | 50.9 | |
| SoftmaxBackbone=IN1K ViT-L, Res.=512^2, Params (M)=304.15, FLOPS (G)=310.60, Peak Mem. (GB)=1.3179, Throughput (imgs/s)=36.47, Segmentation Head=Mask2former2026.03 | 50.83 | |
| SigLIPv2 + SpatialBoostevaluation_protocol=multi-scale evaluation2026.03 | 50.8 | |
| VWFormerYear=2024, Backbone=MiT-B4, Params. (M)=64.0, GFLOPs=79.92024.11 | 50.8 | |
| FeedFormerYear=2023, Backbone=MiT-B4, Params. (M)=65.0, GFLOPs=63.82024.11 | 50.7 | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=642026.05 | 50.69 | |
| dino.txtevaluation_protocol=linear probing2026.03 | 50.6 | |
| DenseClip (ViT-B)Backbone=ViT-B2026.04 | 50.6 | |
| CLUSTSEGYear=2023, Backbone=ResNet-502024.11 | 50.5 | |
| MetaSegYear=2024, Backbone=MiT-B4, Params. (M)=63.6, GFLOPs=55.52024.11 | 50.5 | |
| MaskCLIPBackbone=ViT-B/16, Pre-training Data=YFCC15M, Protocol=supervised finetuning2026.05 | 50.5 | |
| MaskCLIP‡Teacher=CLIP-B, Params=86M, GFLOPs=17.45, Epochs=252026.06 | 50.5 | |
| AM-RADIOv2.5-B/16Protocol=Linear probing, Backbone=B/162026.05 | 50.37 | |
| TSVBase Model=DINOv22026.04 | 50.3 |