Open-Vocabulary Segmentation on Cityscapes
69.7mIoUSegEarth-OV3
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SegEarth-OV3Size=PE-L+/142025.12 | 69.7 | — | |
| OpenSDTraining Data=COCO, Backbone=Swin-L2023.12 | 52.2 | — | |
| X-DecoderTraining Data=COCO+ITP, Backbone=Swin-L2023.12 | 52 | — | |
| CorrCLIPSize=ViT-L/14, Training Protocol=Training-free2025.12 | 51.1 | — | |
| OpenSDTraining Data=COCO, Backbone=Swin-T2023.12 | 51 | — | |
| OpenSDTraining Data=COCO+O365, Backbone=Swin-T2023.12 | 51 | — | |
| OpenSDTraining Data=COCO, Backbone=R502023.12 | 50.2 | — | |
| OpenSDTraining Data=COCO+O365, Backbone=R502023.12 | 50.1 | — | |
| CorrCLIPSize=ViT-H/14, Training Protocol=Training-free2025.12 | 49.9 | — | |
| CorrCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 49.4 | — | |
| OpenSeeDTraining Data=COCO, Backbone=Swin-L2023.12 | 47.6 | — | |
| TridentSize=ViT-H/14, Training Protocol=Training-free2025.12 | 47.6 | — | |
| X-DecoderTraining Data=COCO+ITP, Backbone=Swin-T2023.12 | 47.3 | — | |
| OpenSeeDTraining Data=COCO+O365, Backbone=Swin-T2023.12 | 46.1 | — | |
| OpenSeeDTraining Data=COCO, Backbone=Swin-T2023.12 | 45.8 | — | |
| TridentSize=ViT-B/16, Training Protocol=Training-free2025.12 | 42.9 | — | |
| ProxyCLIPSize=ViT-H/14, Training Protocol=Training-free2025.12 | 42 | — | |
| SC-CLIPSize=ViT-L/14, Training Protocol=Training-free2025.12 | 41.3 | — | |
| SC-CLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 41 | — | |
| ProxyCLIPSize=ViT-L/14, Training Protocol=Training-free2025.12 | 40.1 | — | |
| CASSSize=ViT-B/16, Training Protocol=Training-free2025.12 | 39.4 | — | |
| NACLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 38.3 | — | |
| LoftUp2026.06 | 38.25 | 59.95 | |
| ProxyCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 38.1 | — | |
| JAFAR2026.06 | 38 | 60.32 | |
| AnyUp2026.06 | 37.56 | 60.03 | |
| RaysUp2026.06 | 37.24 | 59.68 | |
| FreeDASize=ViT-L/14, Training Protocol=Training-free2025.12 | 36.7 | — | |
| ResCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 35.9 | — | |
| FeatUp2026.06 | 35.23 | 57.44 | |
| Bilinear2026.06 | 34.91 | 58.52 | |
| ResCLIPSize=ViT-L/14, Training Protocol=Training-free2025.12 | 33.7 | — | |
| SCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 32.2 | — | |
| CLIP-DINOiserSize=ViT-B/16, Training Protocol=Training-based2025.12 | 31.7 | — | |
| INSIGHT_thresParadigm=Use Frozen CLIP with Interpretability, Background Prompt=No2026.01 | 31.5 | — | |
| CLIP-DINOiserParadigm=Use Frozen CLIP, Background Prompt=No, Protocol=Rerun (*)2026.01 | 31.1 | — | |
| ClearCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 30 | — | |
| CoDeSize=ViT-B/16, Training Protocol=Training-based2025.12 | 28.9 | — | |
| INSIGHTParadigm=Use Frozen CLIP with Interpretability, Background Prompt=No2026.01 | 27.8 | — | |
| LaVGSize=ViT-B/16, Training Protocol=Training-free2025.12 | 26.2 | — | |
| JAFARUpsampling=JAFAR, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 25.26 | 61.73 | |
| MaskCLIPParadigm=Use Frozen CLIP, Background Prompt=No2026.01 | 25 | — | |
| FeatUpUpsampling=FeatUp, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 24.76 | 60.11 | |
| TCLSize=ViT-B/16, Training Protocol=Training-based2025.12 | 24 | — | |
| TCLParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 23.1 | — | |
| MaskCLIPParadigm=Use Frozen CLIP, Extra Backbones=ref[9], Background Prompt=No2026.01 | 23 | — | |
| Large Image (×2)Upsampling=Large Image (×2), Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 21.91 | 52.22 | |
| BilinearUpsampling=Bilinear, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 21.56 | 53.21 | |
| ReCoParadigm=Build prototypes per class, Background Prompt=No2026.01 | 21.1 | — | |
| NearestUpsampling=Nearest, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 19.66 | 50.27 | |
| MaskCLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 12.6 | — | |
| CLIP-DIYParadigm=Use Frozen CLIP, Extra Backbones=DINO, Background Prompt=No2026.01 | 11.6 | — | |
| GroupViTParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 11.1 | — | |
| SegCLIPParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 11 | — | |
| CLIPSize=ViT-B/16, Training Protocol=Training-free2025.12 | 5 | — |