Semantic Segmentation on PASCAL Context (mIoU, fwIoU, PACC)
83.43mIoURADIOv2.5
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RADIOv2.5Backbone=ViT-H/162024.12 | 83.43 | — | — | |
| RADIOv2.5Backbone=ViT-L/162024.12 | 82.87 | — | — | |
| AM-RADIOBackbone=ViT-H/162024.12 | 82.78 | — | — | |
| DINOv2Backbone=ViT-g/142024.12 | 82.47 | — | — | |
| DINOv2Backbone=ViT-L/142024.12 | 81.97 | — | — | |
| UNICBackbone=ViT-L/142024.12 | 81.82 | — | — | |
| RADIOv2.5Backbone=ViT-B/162024.12 | 81.75 | — | — | |
| DINOv2Backbone=ViT-B/162024.12 | 81.68 | — | — | |
| MLoREBackbone=Custom2024.12 | 81.41 | — | — | |
| UNICBackbone=ViT-B/162024.12 | 75.9 | — | — | |
| TheiaBackbone=ViT-B/162024.12 | 69.84 | — | — | |
| DINOv3-B/16Protocol=Linear probing, Backbone=B/162026.05 | 57.74 | — | — | |
| SPARKzero-shot=true, training-free=true, features=diffusion2026.01 | 57.7 | — | — | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=642026.05 | 57.46 | — | — | |
| AM-RADIOv2.5-B/16Protocol=Linear probing, Backbone=B/162026.05 | 56.72 | — | — | |
| DiffCutzero-shot=true, training-free=true, features=diffusion2026.01 | 56.5 | — | — | |
| OursBackbone=SDv1.4, Resolution=512 x 5122026.06 | 55.4 | — | — | |
| DINOv2-B/14Protocol=Linear probing, Backbone=B/142026.05 | 55.11 | — | — | |
| DINOv2-reg-B/14Protocol=Linear probing, Backbone=B/142026.05 | 55.02 | — | — | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=82026.05 | 53.3 | — | — | |
| Seg4Diffzero-shot=true, training-free=true, features=diffusion2026.01 | 52.6 | — | — | |
| POMPSource Dataset=Standard COCO Stuff2023.04 | 51.1 | 65.4 | 76.1 | |
| ZSSegSource Dataset=Standard COCO Stuff2023.04 | 50.8 | 64.1 | 75.7 | |
| DiffSegzero-shot=true, training-free=true, features=diffusion2026.01 | 48.8 | — | — | |
| DiffSegBackbone=SDv1.4, Resolution=512 x 5122026.06 | 47.6 | — | — | |
| SigLIP 2-B/16Protocol=Linear probing, Backbone=B/162026.05 | 47.51 | — | — | |
| CLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 45.99 | — | — | |
| DFNCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 44.56 | — | — | |
| OpenCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 44.06 | — | — | |
| MaskCutzero-shot=true, training-free=true, features=diffusion2026.01 | 43.4 | — | — | |
| MaskCut2026.06 | 43.4 | — | — | |
| DiffCutBackbone=SDv1.4, Resolution=512 x 5122026.06 | 40.8 | — | — | |
| DINOv3 + Recursive-NCut2026.06 | 37.2 | — | — | |
| ProxyCLIPBackbone=OpenCLIP-ViT-H/14, Training Strategy=Zero-shot2024.08 | 35.4 | — | — | |
| ProxyCLIPBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 35.3 | — | — | |
| ProxyCLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 34.5 | — | — | |
| GEMBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 32.6 | — | — | |
| CLIP-DINOiserTraining Strategy=Training-based2024.08 | 32.4 | — | — | |
| ProMerge2026.06 | 31.8 | — | — | |
| SCLIPBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 30.4 | — | — | |
| CLIPSurgeryBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 29.3 | — | — | |
| SAM-CLIPTraining Strategy=Training-based2024.08 | 29.2 | — | — | |
| CuVLERUA (Unsupervised Adaptation)=true2026.06 | 25.6 | — | — | |
| SegCLIPTraining Strategy=Training-based2024.08 | 24.7 | — | — | |
| TCLTraining Strategy=Training-based2024.08 | 24.3 | — | — | |
| CoCuTraining Strategy=Training-based2024.08 | 23.6 | — | — | |
| MaskCLIP+Backbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 23.6 | — | — | |
| MaskCLIPzero-shot=true, training-free=true, features=diffusion2026.01 | 23.6 | — | — | |
| MaskCLIPLD (Language Dependency)=true2026.06 | 23.6 | — | — | |
| SCLIPBackbone=OpenCLIP-ViT-H/14, Training Strategy=Zero-shot2024.08 | 23.5 | — | — | |
| ViewCoTraining Strategy=Training-based2024.08 | 23 | — | — | |
| SCLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 22.3 | — | — | |
| OVSegmentorTraining Strategy=Training-based2024.08 | 20.4 | — | — | |
| ReCo+Backbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 19.9 | — | — | |
| ReCOzero-shot=true, training-free=true, features=diffusion2026.01 | 19.9 | — | — | |
| ReCOLD (Language Dependency)=true, AX (Auxiliary image requirement)=true2026.06 | 19.9 | — | — | |
| GroupViTTraining Strategy=Training-based2024.08 | 18.7 | — | — | |
| MaskCLIPBackbone=OpenCLIP-ViT-H/14, Training Strategy=Zero-shot2024.08 | 13.3 | — | — | |
| MaskCLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 11.7 | — | — | |
| CLIPBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 8.4 | — | — | |
| OpenCLIPBackbone=OpenCLIP-ViT-H/14, Training Strategy=Zero-shot2024.08 | 5 | — | — | |
| CLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 4.1 | — | — | |
| FAVisual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 22.5 | — | |
| FA (20M)Training Data=20M high-concept-coverage curated image-text pairs, Visual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 24.6 | — | |
| LinearRSVisual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 8.1 | — | |
| MLPRSVisual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 8.2 | — | |
| PALVisual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 25.5 | — | |
| SAILVisual Backbone=DINOv2 ViT-L, Text Encoder=RoBERTa-Large, Evaluation Protocol=Zero-shot2026.05 | — | 21.1 | — |