Open-Vocabulary Segmentation on Vaihingen (test)
63mIoUCONCEPTBANK
Evaluation Results
| Method | Links | |
|---|---|---|
| CONCEPTBANKBackbone (Size)=SAM3 (PE-L+/14), Training Paradigm=Parameter-free, Evaluation Protocol=with background2026.02 | 63 | |
| SegEarth-OV3Backbone (Size)=SAM3 (PE-L+/14), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=arXiv’252026.02 | 60.8 | |
| SkySense-OBackbone (Size)=CLIP (ViT-L/14@512px), Training Paradigm=Training-based, Evaluation Protocol=with background, Pub. & Year=CVPR’252026.02 | 51.6 | |
| SAM3Backbone (Size)=SAM3 (PE-L+/14), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ICLR’262026.02 | 49.1 | |
| ProxyCLIPBackbone (Size)=CLIP + DINOv2 (ViT-B/14), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ECCV’242026.02 | 47.5 | |
| CAFe-DINORS Training=No, Background Inclusion=Yes, Trained on target dataset=false2026.05 | 47.1 | |
| CorrCLIPBackbone (Size)=CLIP + DINO (ViT-B/8) + SAM2 (Hiera-L), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ICCV’252026.02 | 47 | |
| RSKT-SegBackbone (Size)=CLIP (ViT-L/14@336px) + RemoteCLIP (ViT-B/16) + DINO (ViT-B/32), Training Paradigm=Training-based, Evaluation Protocol=with background, Pub. & Year=AAAI’262026.02 | 42.7 | |
| Cat-SegBackbone (Size)=CLIP (ViT-L/14@336px), Training Paradigm=Training-based, Evaluation Protocol=with background, Pub. & Year=CVPR’242026.02 | 42.3 | |
| SegEarth-OVBackbone (Size)=CLIP (ViT-B/16), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=CVPR’252026.02 | 40 | |
| SANBackbone (Size)=CLIP (ViT-L/14@336px), Training Paradigm=Training-based, Evaluation Protocol=with background, Pub. & Year=CVPR’232026.02 | 39.2 | |
| GEMBackbone (Size)=CLIP (ViT-B/16), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=CVPR’242026.02 | 36.4 | |
| SCLIPBackbone (Size)=CLIP (ViT-B/16), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ECCV’242026.02 | 35.9 | |
| MaskCLIPBackbone (Size)=CLIP (ViT-B/16), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ECCV’222026.02 | 29.9 | |
| GSNetRS Training=Yes, Background Inclusion=Yes, Trained on target dataset=false2026.05 | 28.2 | |
| SegEarth-OVRS Training=Self-Supervised, Background Inclusion=Yes, Trained on target dataset=false2026.05 | 24.3 | |
| OVRSRS Training=Yes, Background Inclusion=Yes, Trained on target dataset=false2026.05 | 19 | |
| DINOv3.txtRS Training=No, Background Inclusion=Yes, Trained on target dataset=false2026.05 | 17.3 | |
| CLIPBackbone (Size)=CLIP (ViT-B/16), Training Paradigm=Parameter-free, Evaluation Protocol=with background, Pub. & Year=ICML’212026.02 | 10.8 |