Semantic Segmentation on COCO Object
64.65mIoUDINOv3-B/16
Evaluation Results
| Method | Links | |
|---|---|---|
| DINOv3-B/16Protocol=Linear probing, Backbone=B/162026.05 | 64.65 | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=642026.05 | 63.58 | |
| AM-RADIOv2.5-B/16Protocol=Linear probing, Backbone=B/162026.05 | 62.94 | |
| DINOv2-reg-B/14Protocol=Linear probing, Backbone=B/142026.05 | 61.77 | |
| DINOv2-B/14Protocol=Linear probing, Backbone=B/142026.05 | 61.01 | |
| VECA-B/16Protocol=Linear probing, Backbone=B/16, Core budget (C)=82026.05 | 58.27 | |
| SigLIP 2-B/16Protocol=Linear probing, Backbone=B/162026.05 | 54.58 | |
| CorrCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 49.4 | |
| CLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 48.76 | |
| CorrCLIPSize=ViT-H/14, Training Approach=Training-free2024.11 | 48.4 | |
| GoCAType=Ours, Backbone=Flux2026.03 | 48.1 | |
| OpenCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 46.27 | |
| DFNCLIP-B/16Protocol=Linear probing, Backbone=B/162026.05 | 46.11 | |
| CLIPtraseSize=ViT-B/16, Training Approach=Training-free2024.11 | 44.8 | |
| GoCAType=Ours, Backbone=SD XL2026.03 | 44.3 | |
| CorrCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 43.7 | |
| CLIPtraseLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 43.6 | |
| CLIPerSize=ViT-L/14, Training Approach=Training-free2024.11 | 43.3 | |
| FluxType=Vanilla2026.03 | 43.3 | |
| DSLO (solve maximum velocity)Logits Model=CLIP ViT-L/142026.04 | 42.9 | |
| SPARKzero-shot=true, training-free=true, features=diffusion2026.01 | 42.7 | |
| DSLO (solve optimal path)Logits Model=CLIP ViT-L/142026.04 | 42.3 | |
| TridentSize=ViT-H/14, Training Approach=Training-free2024.11 | 42.2 | |
| TridentSize=ViT-B/16, Training Approach=Training-free2024.11 | 41.1 | |
| SC-CLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 40.5 | |
| SC-CLIPLogits Model=CLIP ViT-L/14, Paradigm=M.M.2026.04 | 40.5 | |
| GoCAType=Ours, Backbone=Pixart-Sigma2026.03 | 39.8 | |
| DSLO (solve maximum velocity)Logits Model=CLIP ViT-B/162026.04 | 39.6 | |
| ProxyCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 39.2 | |
| ProxyCLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 39.2 | |
| GoCAType=Ours, Backbone=SD v1.52026.03 | 39.2 | |
| CLIPerSize=ViT-B/16, Training Approach=Training-free2024.11 | 39 | |
| DSLO (solve optimal path)Logits Model=CLIP ViT-B/162026.04 | 38.9 | |
| ProxyCLIPSize=ViT-H/14, Training Approach=Training-free2024.11 | 38.6 | |
| ProxyCLIPBackbone=OpenCLIP-ViT-H/14, Training Strategy=Zero-shot2024.08 | 38.6 | |
| Seg4Diffzero-shot=true, training-free=true, features=diffusion2026.01 | 38.5 | |
| DiffSegmentorType=Pre-Trained DM2026.03 | 37.9 | |
| CASSSize=ViT-B/16, Training Approach=Training-free2024.11 | 37.8 | |
| CASSLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 37.8 | |
| SC-CLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 37.7 | |
| SC-CLIPLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 37.7 | |
| ProxyCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 37.5 | |
| ProxyCLIPBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 37.5 | |
| FreeDASize=ViT-L/14, Training Approach=Training-free2024.11 | 37.4 | |
| ProxyCLIP‡Logits Model=CLIP ViT-L/14, Paradigm=M.M.2026.04 | 37.4 | |
| SD XLType=Vanilla2026.03 | 37.2 | |
| SD v1.5Type=Baseline2026.03 | 36.9 | |
| CaRSize=ViT-L/14, Training Approach=Training-free2024.11 | 36.6 | |
| NACLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 36.2 | |
| ProxyCLIP‡Logits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 36.2 | |
| PnP-OVSSLogits Model=CLIP ViT-L/14, Paradigm=M.M.2026.04 | 36.2 | |
| DiffCutTraining Protocol=Training-free, LD=false2024.06 | 36 | |
| ResCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 35 | |
| CLIP-DINOiserTraining Strategy=Training-based2024.08 | 35 | |
| CLIP-DINOiserSize=ViT-B/16, Training Approach=Training-based2024.11 | 34.8 | |
| CLIP-DINOiser†Logits Model=CLIP ViT-B/16, Paradigm=I.T.2026.04 | 34.8 | |
| FTTMType=Pre-Trained DM, Note=Reproduced2026.03 | 34.6 | |
| LaVGSize=ViT-B/16, Training Approach=Training-free2024.11 | 34.2 | |
| LaVGLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 34.2 | |
| DiffCutzero-shot=true, training-free=true, features=diffusion2026.01 | 34.1 | |
| Pixart-SigmaType=Vanilla2026.03 | 33.4 | |
| NACLIPLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 33.2 | |
| ClearCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 33 | |
| ClearCLIPLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 33 | |
| GEMLogits Model=CLIP ViT-B/16, Paradigm=I.T.2026.04 | 32.9 | |
| ResCLIPSize=ViT-L/14, Training Approach=Training-free2024.11 | 32.5 | |
| CoDeSize=ViT-B/16, Training Approach=Training-based2024.11 | 32.3 | |
| SD v1.5Type=Vanilla2026.03 | 32.3 | |
| CLIPpyTraining Protocol=Extra-Training, LD=true2024.06 | 32 | |
| CLIPpyArch=ViT, Dataset=HQITP-134M, SSP=true2022.10 | 32 | |
| TCLSize=ViT-B/16, Training Approach=Training-based2024.11 | 31.6 | |
| TCLTraining Protocol=Extra-Training, LD=true2024.06 | 31.6 | |
| TCLLogits Model=CLIP ViT-B/16, Paradigm=I.T.2026.04 | 31.6 | |
| DiffCutBackbone=SDv1.4, Resolution=512 x 5122026.06 | 31.2 | |
| CLIP-DIYTraining Protocol=Training-free, LD=true2024.06 | 31 | |
| FreeSeg-DiffTraining Protocol=Training-free, LD=false2024.06 | 31 | |
| OursBackbone=SDv1.4, Resolution=512 x 5122026.06 | 30.8 | |
| SCLIPSize=ViT-B/16, Training Approach=Training-free2024.11 | 30.5 | |
| SCLIPBackbone=CLIP-ViT-B/16, Training Strategy=Zero-shot2024.08 | 30.5 | |
| SCLIPLogits Model=CLIP ViT-B/16, Paradigm=M.M.2026.04 | 30.5 | |
| TCLTraining Strategy=Training-based2024.08 | 30.4 | |
| ProMerge2026.06 | 30.2 | |
| MaskCutzero-shot=true, training-free=true, features=diffusion2026.01 | 30.1 | |
| MaskCut2026.06 | 30.1 | |
| NACLIPLogits Model=CLIP ViT-L/14, Paradigm=M.M.2026.04 | 29.9 | |
| CLIPSurgeryLogits Model=CLIP ViT-B/16, Paradigm=I.T.2026.04 | 29.7 | |
| ClearCLIPLogits Model=CLIP ViT-L/14, Paradigm=M.M.2026.04 | 28.6 | |
| CLIPpyArch=ViT, Dataset=CC-12M, SSP=true2022.10 | 28.5 | |
| GEMLogits Model=CLIP ViT-L/14, Paradigm=I.T.2026.04 | 28.3 | |
| CLIPSurgeryLogits Model=CLIP ViT-L/14, Paradigm=I.T.2026.04 | 28.1 | |
| GroupViTTraining Strategy=Training-based2024.08 | 27.5 | |
| GroupViTLogits Model=CLIP ViT-B/16, Paradigm=I.T.2026.04 | 27.5 | |
| DINOv3 + Recursive-NCut2026.06 | 27.5 | |
| SegCLIPTraining Protocol=Extra-Training, LD=false2024.06 | 26.5 | |
| SegCLIPTraining Strategy=Training-based2024.08 | 26.5 | |
| EVACLIP (+LazyStrike)Backbone=ViT-B/16, Zero-shot=true2026.02 | 26.2 | |
| OVSegmentorTraining Protocol=Extra-Training, LD=true2024.06 | 25.1 | |
| OVSArch=ViT, Dataset=CC-12M, SSP=true2022.10 | 25.1 | |
| OVSegmentorTraining Strategy=Training-based2024.08 | 25.1 | |
| SCLIPBackbone=CLIP-ViT-L/14, Training Strategy=Zero-shot2024.08 | 25 |