Semantic Segmentation on LVIS 92^i
47.2mIoUINSID3
Evaluation Results
| Method | Links | |
|---|---|---|
| INSID3Encoder=DINOv3, #Param=304 M, Training Protocol=Training free: Unsupervised pre-training, Shot count=52026.03 | 47.2 | |
| SegICEncoder=DINOv2, #Param=310 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 44.6 | |
| GF-SAMEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 44.2 | |
| GF-SAM† + our debiasEncoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 43.6 | |
| GF-SAM†Encoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 42.8 | |
| INSID3Encoder=DINOv3, #Param=304 M, Training Protocol=Training free, Supervision Type=Unsupervised pre-training2026.03 | 41.8 | |
| MatcherShots=few-shot, Setting=In-context2024.10 | 40 | |
| MatcherEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free: Mask-supervised pre-training, Shot count=52026.03 | 40 | |
| SegIC (COCO)Encoder=DINOv2, #Param=310 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision, dataset split=COCO2026.03 | 35.7 | |
| SINEEncoder=DINOv2, #Param=373 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 35.5 | |
| DiffewSShots=few-shot, Setting=In-context2024.10 | 35.4 | |
| DiffewSEncoder=Stable Diffusion, #Param=890 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 35.4 | |
| GF-SAMEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 35.2 | |
| GF-SAM† + debiasEncoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training, Debias=True2026.03 | 34.6 | |
| MatcherShots=one-shot, Setting=In-context2024.10 | 33 | |
| MatcherEncoder=DINOv2 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 33 | |
| GF-SAM†Encoder=DINOv3 + SAM, #Param=945 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 31.8 | |
| DiffewSShots=one-shot, Setting=In-context2024.10 | 31.4 | |
| DiffewSEncoder=Stable Diffusion, #Param=890 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 31.4 | |
| SINEEncoder=DINOv2, #Param=373 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 31.2 | |
| SegGPTShots=few-shot, Setting=In-context2024.10 | 25.4 | |
| SegGPTEncoder=ViT, #Param=354 M, Training Protocol=Task-specific fine-tuning: Semantic + mask supervision, Shot count=52026.03 | 25.4 | |
| HSNetShots=few-shot, Setting=In-context2024.10 | 22.9 | |
| VATShots=few-shot, Setting=In-context2024.10 | 22.7 | |
| SegGPTShots=one-shot, Setting=In-context2024.10 | 18.6 | |
| SegGPTEncoder=ViT, #Param=354 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 18.6 | |
| VATShots=one-shot, Setting=In-context2024.10 | 18.5 | |
| PerSAM-FShots=one-shot, Setting=In-context2024.10 | 18.4 | |
| HSNetShots=one-shot, Setting=In-context2024.10 | 17.4 | |
| PerSAMShots=one-shot, Setting=In-context2024.10 | 15.6 | |
| PerSAM-Fone-shot=true2025.12 | 12.3 | |
| PerSAMone-shot=true2025.12 | 11.5 | |
| PerSAMEncoder=SAM, #Param=640 M, Training Protocol=Training free, Supervision Type=Mask-supervised pre-training2026.03 | 11.5 | |
| PainterShots=few-shot, Setting=In-context2024.10 | 10.9 | |
| PainterShots=one-shot, Setting=In-context2024.10 | 10.5 | |
| Painterone-shot=true2025.12 | 10.5 | |
| PainterEncoder=ViT, #Param=354 M, Training Protocol=Task-specific fine-tuning, Supervision Type=Semantic + mask supervision2026.03 | 10.5 | |
| OmniSegNetone-shot=true2025.12 | 9.2 |