Semantic Segmentation on Cityscapes (Semseg mIoU)
93.95mIoUB^3-Net
Evaluation Results
| Method | Links | |
|---|---|---|
| B^3-Net2026.05 | 93.95 | |
| BridgeNet2026.05 | 93.72 | |
| MTMamba++2026.05 | 91.11 | |
| MTMamba2026.05 | 90.77 | |
| InvPT2026.05 | 89.86 | |
| MTAN2026.05 | 88.9 | |
| PAD-Net2026.05 | 85.59 | |
| TaskPrompter2026.05 | 85.26 | |
| SwinMTL2026.05 | 85.04 | |
| Lorentz (SegFormer-B3)Backbone=B32026.04 | 83.7 | |
| EoMTBackbone=ViT-Large, FPS A100=17.6, FPS Orin=3.6, GFLOPs=1721.3, Resolution=1024 × 10242026.05 | 83.6 | |
| SCASegBackbone=MiT-B5, Params. (M)=92.4, GFLOPs=1173.02024.11 | 83.5 | |
| NaLaFormer-SPARA=25M, F=206G2025.06 | 83.5 | |
| SCASegBackbone=MiT-B4, Params. (M)=71.8, GFLOPs=953.02024.11 | 83.2 | |
| ContrastiveSegYear=2021, Backbone=OCR2024.11 | 83.2 | |
| U-MixFormerYear=2025, Backbone=MiT-B5, Params. (M)=93.0, GFLOPs=1171.02024.11 | 83.1 | |
| SCASegBackbone=MSCAN-B, Params. (M)=36.5, GFLOPs=261.02024.11 | 83 | |
| SCASegBackbone=MiT-B3, Params. (M)=55.1, GFLOPs=675.02024.11 | 83 | |
| TokenMaskBackbone=ViT-Large, FPS A100=21.9, FPS Orin=4.2, GFLOPs=1251.6, Resolution=1024 × 10242026.05 | 82.9 | |
| VWFormerYear=2024, Backbone=MiT-B5, Params. (M)=84.6, GFLOPs=1213.02024.11 | 82.8 | |
| MetaSegYear=2024, Backbone=MSCAN-B, Params. (M)=29.6, GFLOPs=251.12024.11 | 82.7 | |
| VWFormerYear=2024, Backbone=MiT-B4, Params. (M)=64.0, GFLOPs=993.02024.11 | 82.7 | |
| FeedFormerYear=2023, Backbone=MiT-B5, Params. (M)=85.6, GFLOPs=1180.02024.11 | 82.7 | |
| SegNeXtYear=2022, Backbone=MSCAN-B, Params. (M)=27.6, GFLOPs=275.72024.11 | 82.6 | |
| FeedFormerYear=2023, Backbone=MiT-B4, Params. (M)=65.0, GFLOPs=960.02024.11 | 82.6 | |
| MetaSegYear=2024, Backbone=MiT-B5, Params. (M)=85.0, GFLOPs=1143.02024.11 | 82.5 | |
| NaLaFormer-TPARA=14M, F=111G2025.06 | 82.5 | |
| VWFormerYear=2024, Backbone=MiT-B3, Params. (M)=47.3, GFLOPs=715.02024.11 | 82.4 | |
| SegFormerYear=2021, Backbone=MiT-B5, Params. (M)=84.7, GFLOPs=1460.42024.11 | 82.4 | |
| SegFormer (B4)Backbone=B42026.04 | 82.3 | |
| VWFormerYear=2024, Backbone=MSCAN-B, Params. (M)=28.3, GFLOPs=302.02024.11 | 82.3 | |
| M2H-MX-L2026.03 | 82.28 | |
| M2H-MX-L2026.05 | 82.28 | |
| SETRYear=2021, Backbone=ViT-Large, Params. (M)=318.32024.11 | 82.2 | |
| MTA-clip (ViT-B)Backbone=ViT-B2026.04 | 82.1 | |
| FeedFormerYear=2023, Backbone=MSCAN-B, Params. (M)=30.5, GFLOPs=269.02024.11 | 82.1 | |
| MetaSegYear=2024, Backbone=MiT-B4, Params. (M)=63.6, GFLOPs=923.02024.11 | 82.1 | |
| EfficientViT-B2PARA=15M, F=74G2025.06 | 82.1 | |
| FeedFormerYear=2023, Backbone=MiT-B3, Params. (M)=48.3, GFLOPs=682.02024.11 | 81.9 | |
| SegFormerYear=2021, Backbone=MiT-B4, Params. (M)=64.1, GFLOPs=1240.62024.11 | 81.9 | |
| MetaSegYear=2024, Backbone=MiT-B3, Params. (M)=47.7, GFLOPs=645.02024.11 | 81.8 | |
| EoMTBackbone=ViT-Base, FPS A100=37.5, FPS Orin=9.6, GFLOPs=628.1, Resolution=1024 × 10242026.05 | 81.8 | |
| SegFormer (B3)Backbone=B32026.04 | 81.7 | |
| SegFormerYear=2021, Backbone=MiT-B3, Params. (M)=47.3, GFLOPs=962.92024.11 | 81.7 | |
| VWFormer-B2PARA=27M, F=415G2025.06 | 81.7 | |
| ContrastiveSegYear=2021, Backbone=HRNetV2-W482024.11 | 81.4 | |
| SegNeXt-SPARA=15M, F=125G2025.06 | 81.3 | |
| DINOv3Number of Parameters=7B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Evaluation Resolution=default2026.07 | 81.1 | |
| DenseClip (R101)Backbone=R1012026.04 | 81 | |
| EoMTBackbone=ViT-Small, FPS A100=78.3, FPS Orin=22.6, GFLOPs=165.4, Resolution=1024 × 10242026.05 | 81 | |
| SegFormer-B2PARA=28M, F=711G2025.06 | 81 | |
| Lorentz (DeepLab-R101)Backbone=R1012026.04 | 80.6 | |
| TokenMaskBackbone=ViT-Base, FPS A100=51.4, FPS Orin=11.9, GFLOPs=356.6, Resolution=1024 × 10242026.05 | 80.4 | |
| VWFormer-B1PARA=14M2025.06 | 80.4 | |
| SupervisedAnnotation=Fine, Architecture=DeepLabV3+ (ResNet-101)2026.04 | 80.2 | |
| DeepLabV32026.04 | 80.1 | |
| TokenMaskBackbone=ViT-Small, FPS A100=104.4, FPS Orin=27.2, GFLOPs=90.4, Resolution=1024 × 10242026.05 | 79.7 | |
| LingBot-Vision ViT-gNumber of Parameters=1B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Evaluation Resolution=default2026.07 | 79.6 | |
| DINOv3 ViT-H+Number of Parameters=0.8B, Patch Size=16, Evaluation Protocol=Linear probing on frozen features, Evaluation Resolution=default2026.07 | 79.5 | |
| MTMamba++2026.03 | 79.13 | |
| MTMamba++2026.05 | 79.13 | |
| SegmenterBackbone=ViT-Large, FPS A100=20.8, FPS Orin=3.4, GFLOPs=2248.3, Resolution=1024 × 10242026.05 | 79.1 | |
| SegFormer-B1PARA=14M, F=244G2025.06 | 78.5 | |
| AM-RADIOv2.5Number of Parameters=1B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Evaluation Resolution=default2026.07 | 78.4 | |
| CTSBackbone=ViT-B/16, FPS (B=32)=80, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=true2026.04 | 78.11 | |
| TeacherBackbone=ResNet-101, Segmentation Framework=DeepLabV32024.11 | 78.07 | |
| MTMamba2026.03 | 78 | |
| MTMamba2026.05 | 78 | |
| Mask2Former-R50GFLOPs=527.4, Params (M)=44.0, FPS=2.42026.04 | 77.5 | |
| ViT-B/16Backbone=ViT-B/16, GFLOPs=348, FPS (B=32)=44, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 77.3 | |
| ToMeBackbone=ViT-B/16, GFLOPs=∼298, FPS (B=32)=46, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 77.3 | |
| DHKDTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 77.13 | |
| DISTTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 77.1 | |
| SegmenterBackbone=ViT-Base, FPS A100=6.8, FPS Orin=9.9, GFLOPs=776.0, Resolution=1024 × 10242026.05 | 77.1 | |
| ALGMBackbone=ViT-S/16, GFLOPs=79, FPS (B=32)=132, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=true2026.04 | 76.9 | |
| CTSBackbone=ViT-S/16, FPS (B=32)=169, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=true2026.04 | 76.6 | |
| SegmenterBackbone=ViT-Small, FPS A100=15.7, FPS Orin=22.1, GFLOPs=285.0, Resolution=1024 × 10242026.05 | 76.6 | |
| ALGMBackbone=ViT-B/16, GFLOPs=∼ 211, FPS (B=32)=62, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=true2026.04 | 76.4 | |
| CIRKDTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 76.38 | |
| ViT-S/16Backbone=ViT-S/16, GFLOPs=116, FPS (B=32)=106, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 76.3 | |
| ToMeBackbone=ViT-S/16, GFLOPs=∼97, FPS (B=32)=105, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 76.3 | |
| MPM (2,5)Backbone=ViT-B/16, GFLOPs=∼ 247, FPS (B=32)=65, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 76.3 | |
| SM-AdaptFormerInput size=1024 × 10242024.08 | 76.25 | |
| SupervisedAnnotation=Fine, Architecture=SegFormer (MiT-B0)2026.04 | 76 | |
| DINOv2Number of Parameters=1B, Patch Size=14, Evaluation Protocol=Linear probing on frozen features, Evaluation Resolution=default2026.07 | 75.6 | |
| IFVDTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 75.59 | |
| CWDTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 75.55 | |
| AdaptFormerInput size=1024 × 10242024.08 | 75.49 | |
| SKDTeacher Backbone=ResNet-101, Student Backbone=ResNet-18, Segmentation Framework=DeepLabV32024.11 | 75.42 | |
| MPM (2,5)Backbone=ViT-S/16, GFLOPs=∼ 81, FPS (B=32)=158, Segmentation Head=Mask Transformer [26], Input Resolution=768x768, Precision=float32, Hardware=H100, Training/Fine-tuning Required=false2026.04 | 75.3 | |
| FixMatch*Sampling=Explicit, Overlap size=7532026.05 | 75.23 | |
| DenseFixMatchSampling=Implicit, Overlap size=7532026.05 | 75.17 | |
| SeSAMAnnotation=Scribble, Architecture=DeepLabV3+ (ResNet-101)2026.04 | 75.1 | |
| DenseFixMatchSampling=Explicit, Overlap size=7532026.05 | 75.1 | |
| DenseFixMatchSampling=Implicit, # labels=14882026.05 | 75.06 | |
| FixMatch*Sampling=Implicit, Overlap size=7532026.05 | 75.04 | |
| FixMatch*Sampling=Implicit, # labels=14882026.05 | 74.91 | |
| FixMatch*Sampling=Explicit, # labels=14882026.05 | 74.91 | |
| DenseFixMatchSampling=Explicit, # labels=14882026.05 | 74.89 | |
| DenseFixMatchSampling=Implicit, # labels=7442026.05 | 74.37 |