Semantic Segmentation on VDD
85.75mIoUSegFormer MiT-B2
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SegFormer MiT-B2Type=Transformer, Param. (M)=27.461, GFLOPs=164.70, log(Param.)=1.439, FPS=2.82, Input image size=1000 x 10002026.02 | 85.75 | 10.38 | 4.38 | — | |
| UperNet SwinLType=Hybrid, Param. (M)=233.962, GFLOPs=825.92, log(Param.)=2.369, FPS=0.08, Input image size=1000 x 10002026.02 | 85.63 | 10.26 | 0.52 | — | |
| UperNet SwinTType=Hybrid, Param. (M)=59.941, GFLOPs=811.59, log(Param.)=1.778, FPS=0.10, Input image size=1000 x 10002026.02 | 84.73 | 9.36 | 0.65 | — | |
| Mask2Former ResNet50Type=Hybrid, Param. (M)=45.517, GFLOPs=133.70, log(Param.)=1.658, FPS=3.48, Input image size=1000 x 10002026.02 | 83.21 | 7.84 | 3.54 | — | |
| SegFormer MiT-B5Type=Transformer, Param. (M)=84.708, GFLOPs=180.53, log(Param.)=1.928, FPS=1.53, Input image size=1000 x 10002026.02 | 82.11 | 6.74 | 1.94 | — | |
| DAS-SKType=CNN, Param. (M)=10.678, GFLOPs=43.52, log(Param.)=1.028, FPS=10.33, Input image size=1000 x 10002026.02 | 79.45 | 4.08 | 9.12 | — | |
| Mask2Former SwinTType=Transformer, Param. (M)=47.439, GFLOPs=45.74, log(Param.)=1.676, FPS=3.04, Input image size=1000 x 10002026.02 | 77.85 | 2.48 | 3.23 | — | |
| SegFormer MiT-B0Type=Transformer, Param. (M)=3.752, GFLOPs=20.89, log(Param.)=0.574, FPS=13.62, Input image size=1000 x 10002026.02 | 75.37 | — | — | — | |
| SegEarth-OV3Training Regime=Training-free2025.12 | 64.5 | — | — | — | |
| OracleTraining Regime=Oracle (fully supervised SegFormer-b0)2025.12 | 62.9 | — | — | — | |
| VIPBackbone=DINOv3 (ViT-L)2026.05 | 54.3 | — | — | — | |
| ConInfer2026.03 | 50.29 | — | — | — | |
| ProxyCLIPTraining Regime=Training-free2025.12 | 47.8 | — | — | — | |
| Pi-SegBackbone=ViT-L, Type=OVRSIS, Training Dataset=DLRSD2026.04 | 47.73 | — | — | 67.1 | |
| CorrCLIPTraining Regime=Training-free2025.12 | 47.3 | — | — | — | |
| SegEarth-OV2026.03 | 46.73 | — | — | — | |
| Trident2026.05 | 45.7 | — | — | — | |
| SegEarth-OVVenue=Ours2024.10 | 45.3 | — | — | — | |
| SegEarth-OVTraining Regime=Training-free2025.12 | 45.3 | — | — | — | |
| SegEarth-OV2026.05 | 45.3 | — | — | — | |
| ProxyCLIP + GLA2026.03 | 45 | — | — | — | |
| GLA-CLIP2026.03 | 45 | — | — | — | |
| ProxyCLIP2026.03 | 44.3 | — | — | — | |
| ProxyCLIP2026.05 | 44.3 | — | — | — | |
| Pi-SegBackbone=ViT-B, Type=OVRSIS, Training Dataset=DLRSD2026.04 | 44.15 | — | — | 63.72 | |
| MaskCLIP*type=baseline2026.03 | 42.54 | — | — | — | |
| RSKT-SegBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 41.22 | — | — | 58.09 | |
| DR-SegBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 41 | — | — | 60.16 | |
| SC-CLIP2026.05 | 41 | — | — | — | |
| RSKT-SegBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 40.55 | — | — | 58.75 | |
| ClearCLIP2026.03 | 40.24 | — | — | — | |
| RSKT-SegTraining Regime=Training on remote sensing segmentation data2025.12 | 39.7 | — | — | — | |
| GEMVenue=CVPR 242024.10 | 39.5 | — | — | — | |
| GEMTraining Regime=Training-free2025.12 | 39.5 | — | — | — | |
| DR-SegBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 39.35 | — | — | 54.79 | |
| ClearCLIPVenue=ECCV'242024.10 | 39.3 | — | — | — | |
| ClearCLIPTraining Regime=Training-free2025.12 | 39.3 | — | — | — | |
| ClearCLIP2026.03 | 39.3 | — | — | — | |
| Cat-SegBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 39.14 | — | — | 55.85 | |
| Cat-SegTraining Regime=Training on remote sensing segmentation data2025.12 | 39.1 | — | — | — | |
| GEM2026.03 | 38.46 | — | — | — | |
| Pi-SegBackbone=ViT-L, Type=OVRSIS, Training Dataset=iSAID2026.04 | 38.24 | — | — | 54.25 | |
| GSNetBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 38.1 | — | — | 56.01 | |
| SCLIPVenue=arXiv'232024.10 | 37.9 | — | — | — | |
| SCLIPTraining Regime=Training-free2025.12 | 37.9 | — | — | — | |
| CLIP-DINOiser2026.03 | 37.7 | — | — | — | |
| CorrCLIP2026.05 | 37.7 | — | — | — | |
| OVRSBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 37.34 | — | — | 55.15 | |
| GSNetBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 37.34 | — | — | 57.04 | |
| GSNetTraining Regime=Training on remote sensing segmentation data2025.12 | 37.3 | — | — | — | |
| OVRSBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 37.23 | — | — | 56.34 | |
| OVRSTraining Regime=Training on remote sensing segmentation data2025.12 | 37.2 | — | — | — | |
| Cat-SegBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 36.18 | — | — | 54.3 | |
| SANBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 35.83 | — | — | 53.25 | |
| SANTraining Regime=Training on remote sensing segmentation data2025.12 | 35.8 | — | — | — | |
| SANBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 34.76 | — | — | 52.42 | |
| DR-SegBackbone=ViT-L, Training Dataset=iSAID2026.04 | 34.75 | — | — | 60.23 | |
| SCLIP2026.03 | 34.26 | — | — | — | |
| RSKT-SegBackbone=ViT-L, Training Dataset=iSAID2026.04 | 33.4 | — | — | 56.5 | |
| MaskCLIPVenue=ECCV'222024.10 | 32.9 | — | — | — | |
| MaskCLIPTraining Regime=Training-free2025.12 | 32.9 | — | — | — | |
| SEDBackbone=ConvNeXt-L, Training Dataset=DLRSD2026.04 | 32.53 | — | — | 51.34 | |
| SEDTraining Regime=Training on remote sensing segmentation data2025.12 | 32.5 | — | — | — | |
| dino.txt2026.05 | 32.3 | — | — | — | |
| DR-SegBackbone=ViT-B, Training Dataset=iSAID2026.04 | 32.2 | — | — | 53.41 | |
| GSNetBackbone=ViT-L, Training Dataset=iSAID2026.04 | 32.07 | — | — | 55.21 | |
| SEDBackbone=ConvNeXt-B, Training Dataset=DLRSD2026.04 | 31.43 | — | — | 50.25 | |
| OVRSBackbone=ViT-L, Training Dataset=iSAID2026.04 | 31.01 | — | — | 55.3 | |
| Cat-SegBackbone=ViT-L, Training Dataset=iSAID2026.04 | 30.16 | — | — | 47.2 | |
| SCANBackbone=ViT-L, Training Dataset=DLRSD2026.04 | 29.24 | — | — | 45.57 | |
| SCANTraining Regime=Training on remote sensing segmentation data2025.12 | 29.2 | — | — | — | |
| MaskCLIP2026.03 | 29.12 | — | — | — | |
| SEDBackbone=ConvNeXt-L, Training Dataset=iSAID2026.04 | 28.5 | — | — | 44.2 | |
| SANBackbone=ViT-L, Training Dataset=iSAID2026.04 | 26.42 | — | — | 42.3 | |
| SCANBackbone=ViT-B, Training Dataset=DLRSD2026.04 | 26.25 | — | — | 43.67 | |
| RSKT-SegBackbone=ViT-B, Training Dataset=iSAID2026.04 | 25.34 | — | — | 47.18 | |
| GSNetBackbone=ViT-B, Training Dataset=iSAID2026.04 | 22.22 | — | — | 42.33 | |
| SANBackbone=ViT-B, Training Dataset=iSAID2026.04 | 22.1 | — | — | 40.25 | |
| OVRSBackbone=ViT-B, Training Dataset=iSAID2026.04 | 21.42 | — | — | 51.25 | |
| SCANBackbone=ViT-L, Training Dataset=iSAID2026.04 | 20.18 | — | — | 36.7 | |
| Cat-SegBackbone=ViT-B, Training Dataset=iSAID2026.04 | 19.62 | — | — | 51.38 | |
| SEDBackbone=ConvNeXt-B, Training Dataset=iSAID2026.04 | 17.32 | — | — | 38.12 | |
| SCANBackbone=ViT-B, Training Dataset=iSAID2026.04 | 15.32 | — | — | 35.1 | |
| CLIPVenue=ICML 212024.10 | 14.2 | — | — | — | |
| CLIPTraining Regime=Training-free2025.12 | 14.2 | — | — | — | |
| CLIP2026.05 | 14.2 | — | — | — | |
| CLIP2026.03 | 11.27 | — | — | — |