Open-Vocabulary Segmentation on ADE20K
47.47mIoUViT-Adapter
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ViT-AdapterEvaluation Protocol=Most similar evaluation, Semantics Level=Full2023.10 | 47.47 | — | — | — | — | |
| LSegEvaluation Protocol=Most similar evaluation, Semantics Level=Full2023.10 | 27.4 | — | — | — | — | |
| LoftUp2026.06 | 21.1 | 43.28 | — | — | — | |
| AnyUp2026.06 | 20.67 | 42.77 | — | — | — | |
| JAFAR2026.06 | 20.57 | 42.76 | — | — | — | |
| RaysUp2026.06 | 20.54 | 42.71 | — | — | — | |
| INSIGHT_thresParadigm=Use Frozen CLIP with Interpretability, Background Prompt=No2026.01 | 20.4 | — | — | — | — | |
| CLIP-DINOiserParadigm=Use Frozen CLIP, Background Prompt=No, Protocol=Rerun (*)2026.01 | 20.1 | — | — | — | — | |
| FeatUp2026.06 | 19.78 | 43.09 | — | — | — | |
| Bilinear2026.06 | 19.6 | 42.32 | — | — | — | |
| INSIGHTParadigm=Use Frozen CLIP with Interpretability, Background Prompt=No2026.01 | 18.2 | — | — | — | — | |
| OVDiffParadigm=Build prototypes per class, Extra Backbones=DINO & SD, Background Prompt=No2026.01 | 14.9 | — | — | — | — | |
| TCLParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 14.9 | — | — | — | — | |
| MaskCLIPParadigm=Use Frozen CLIP, Extra Backbones=ref[9], Background Prompt=No2026.01 | 14.9 | — | — | — | — | |
| MaskCLIPParadigm=Use Frozen CLIP, Background Prompt=No2026.01 | 14.3 | — | — | — | — | |
| JAFARUpsampling=JAFAR, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 13.61 | 33.28 | — | — | — | |
| CLIPpyParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 13.5 | — | — | — | — | |
| FeatUpUpsampling=FeatUp, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 13.03 | 33.28 | — | — | — | |
| ReCoParadigm=Build prototypes per class, Background Prompt=No2026.01 | 11.2 | — | — | — | — | |
| BilinearUpsampling=Bilinear, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 11.03 | 27.78 | — | — | — | |
| ZSSegEvaluation Protocol=Most similar evaluation, Semantics Level=Full2023.10 | 9.93 | — | — | — | — | |
| CLIP-DIYParadigm=Use Frozen CLIP, Extra Backbones=DINO, Background Prompt=No2026.01 | 9.9 | — | — | — | — | |
| NearestUpsampling=Nearest, Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 9.33 | 24.65 | — | — | — | |
| GroupViTParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 9.2 | — | — | — | — | |
| SegCLIPParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 8.7 | — | — | — | — | |
| Large Image (×2)Upsampling=Large Image (×2), Zero-shot=true, Backbone=CLIP-ViT-B/16, Framework=MaskCLIP [9]2025.06 | 8.08 | 24.94 | — | — | — | |
| X-DecoderEvaluation Protocol=Most similar evaluation, Semantics Level=Full2023.10 | 6.48 | — | — | — | — | |
| OVSegmentorParadigm=Text/image alignment training with captions, Background Prompt=No2026.01 | 5.6 | — | — | — | — | |
| GGNRanking=OPA2022.04 | — | — | 18.3 | 7.9 | — | |
| GGNRanking=OPA + OOLN2022.04 | — | — | 21 | 9.7 | — | |
| GGNRanking=GGN, pre-training=pseudo-GT pre-training2022.04 | — | — | 21.5 | 9.3 | — | |
| GGN2023.03 | — | — | — | — | 21 | |
| Mask R-CNN2022.04 | — | — | 14.7 | 6.4 | — | |
| ODISE2023.03 | — | — | — | — | 30.3 | |
| Selective Search2022.04 | — | — | 3.8 | — | — |