Instance Segmentation on ScanNet (val)
44.2mAPOccuSeg
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OccuSeg2019.11 | 44.2 | 60.7 | 71.9 | — | — | |
| Sonata (full)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 42.4 | 63.9 | 79.2 | — | — | |
| SonataReal Pretraining Data=18k, Synth Pretraining Data=121k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 42.4 | — | — | — | — | |
| PTv3 + PPTBackbone=PTv3, Pre-training=PPT, Mode=Fine-tuned, Instance Segmentation Framework=PointGroup, Params.=46.3M2023.08 | 42.1 | 63.5 | 78.9 | — | — | |
| PPT (sup.)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 42.1 | 63.5 | 78.9 | — | — | |
| PPTReal Pretraining Data=1k, Synth Pretraining Data=21k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 42.1 | — | — | — | — | |
| LAM3C*VGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | 41.7 | — | — | — | — | |
| PointINSProtocol=Fine-tuning2026.03 | 41.5 | 63.7 | 78.4 | — | — | |
| MSC (full)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 41.1 | 62.9 | 78.4 | — | — | |
| MSCReal Pretraining Data=7k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 41.1 | — | — | — | — | |
| PTv3Backbone=PTv3, Pre-training=None, Instance Segmentation Framework=PointGroup, Params.=46.2M2023.08 | 40.9 | 61.7 | 77.5 | — | — | |
| PointGroupParams=124.8M, Learn.=124.8M, Pct.=100%2025.03 | 40.9 | 61.7 | 77.5 | — | — | |
| PTv3Evaluation protocol=supervised2026.03 | 40.9 | 61.7 | 77.5 | — | — | |
| PTv3 (sup.)Protocol=Supervised2026.03 | 40.9 | 61.7 | 77.5 | — | — | |
| Sonata (dec.)Params=124.8M, Learn.=16.3M, Pct.=13%, Evaluation Protocol=decoder probing2025.03 | 40.8 | 62.8 | 76.8 | — | — | |
| SparseUNet + PPTBackbone=SparseUNet, Pre-training=PPT, Mode=Fine-tuned, Instance Segmentation Framework=PointGroup, Params.=41.0M2023.08 | 40.7 | 62 | 76.9 | — | — | |
| LAM3C*Real Pretraining Data=15k, VGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | 40.6 | — | — | — | — | |
| DOSProtocol=Fine-tuning2026.03 | 40.5 | 62 | 77.3 | — | — | |
| Sonata (all real)Real Pretraining Data=15k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 40.3 | — | — | — | — | |
| PointINSProtocol=Decoder2026.03 | 40.2 | 62.5 | 77.2 | — | — | |
| SegContrastProtocol=Fine-tuning2026.03 | 39.8 | 60.5 | 75.2 | — | — | |
| LAM3CVGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 39.7 | — | — | — | — | |
| SonataProtocol=Fine-tuning2026.03 | 39.5 | 61.1 | 76.7 | — | — | |
| SparseUNet + MSCBackbone=SparseUNet, Pre-training=MSC, Instance Segmentation Framework=PointGroup, Params.=39.2M2023.08 | 39.3 | 59.6 | 74.7 | — | — | |
| LAM3CVGPC Pretraining Data=16k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 38.9 | — | — | — | — | |
| DOSProtocol=Decoder2026.03 | 38.9 | 60.2 | 75.5 | — | — | |
| PSAProtocol=Fine-tuning2026.03 | 38.8 | 59.7 | 74.3 | — | — | |
| SonataProtocol=Decoder2026.03 | 37.1 | 57.8 | 74.2 | — | — | |
| NOMAEProtocol=Fine-tuning2026.03 | 37 | 59.5 | 76.4 | — | — | |
| SparseUNetBackbone=SparseUNet, Pre-training=None, Instance Segmentation Framework=PointGroup, Params.=39.2M2023.08 | 36 | 56.9 | 72.8 | — | — | |
| 3D-MPA2019.11 | 35.3 | 59.1 | 72.4 | — | — | |
| PointGroup2019.11 | 34.8 | 56.9 | 71.3 | — | — | |
| PointINS*Evaluation protocol=linear probing2026.03 | 34.6 | 57.6 | 76.3 | — | — | |
| ProposedBackbone=MinkowskiNet2019.11 | 33 | 57.1 | 73.8 | — | — | |
| DOS*Evaluation protocol=linear probing2026.03 | 32.5 | 54.6 | 70.9 | — | — | |
| PointINSProtocol=Linear Probing2026.03 | 32.1 | 55.2 | 73.6 | — | — | |
| LAM3C*Real Pretraining Data=15k, VGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | 30.9 | — | — | — | — | |
| Sonata (lin.)Params=124.8M, Learn.=<0.2M, Pct.=<0.2%, Evaluation Protocol=linear probing2025.03 | 30.7 | 53.9 | 72.6 | — | — | |
| Sonata*Evaluation protocol=linear probing2026.03 | 30.7 | 53.9 | 72.6 | — | — | |
| SegContrastProtocol=Decoder2026.03 | 30.5 | 50.3 | 66.9 | — | — | |
| Sonata (ScanNet)Real Pretraining Data=1k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 29.8 | — | — | — | — | |
| PSAProtocol=Decoder2026.03 | 28.9 | 47.3 | 64.5 | — | — | |
| DOSProtocol=Linear Probing2026.03 | 28.7 | 49.8 | 68.7 | — | — | |
| LAM3C*VGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | 28.6 | — | — | — | — | |
| NOMAEProtocol=Decoder2026.03 | 28.5 | 49.3 | 68 | — | — | |
| Learnable marginBackbone=MinkowskiNet2019.11 | 28.1 | 50.1 | 70.1 | — | — | |
| Sonata (all real)Real Pretraining Data=15k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | 28 | — | — | — | — | |
| PTv3Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | 26.9 | — | — | — | — | |
| LAM3CVGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | 25.1 | — | — | — | — | |
| SonataProtocol=Linear Probing2026.03 | 25 | 46.1 | 64.6 | — | — | |
| Sonata (ScanNet)Real Pretraining Data=1k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | 23.9 | — | — | — | — | |
| MTML2019.11 | 20.3 | 40.2 | 55.4 | — | — | |
| LAM3CVGPC Pretraining Data=16k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | 19.1 | — | — | — | — | |
| FreePointunsupervised=class-agnostic2023.05 | 18.9 | 36.4 | — | — | — | |
| PSAProtocol=Linear Probing2026.03 | 9.7 | 20.4 | 41.9 | — | — | |
| NOMAEProtocol=Linear Probing2026.03 | 9.5 | 20 | 42 | — | — | |
| SegContrastProtocol=Linear Probing2026.03 | 6.4 | 13.7 | 30.6 | — | — | |
| DBSCANunsupervised=class-agnostic2023.05 | 3.3 | 3.6 | — | — | — | |
| MSC (lin.)Params=124.8M, Learn.=<0.2M, Pct.=<0.2%, Evaluation Protocol=linear probing2025.03 | 2.3 | 5.3 | 13.3 | — | — | |
| Nunes et al.unsupervised=class-agnostic2023.05 | 2.1 | 6.9 | — | — | — | |
| HDBSCANunsupervised=class-agnostic2023.05 | 1.9 | 5.4 | — | — | — | |
| PTv3Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | 0.2 | — | — | — | — | |
| CityScapes-trained modelTraining Dataset=CityScapes2021.02 | — | — | — | 0 | — | |
| COCO-trained modelTraining Dataset=COCO2021.02 | — | — | — | 5.2 | — | |
| CSCBackbone=SparseUNet, Pre-training=Contrastive Scene Contexts, Downstream Framework=PointGroup2023.03 | — | 59.4 | — | — | — | |
| CSCBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | — | 3.5 | — | — | — | |
| EMSAFormerBackbone=SwinV2-T-128, Augmentation Strategy=Multi-Aug2023.06 | — | — | — | — | 66.71 | |
| EMSAFormerBackbone=SwinV2-T-128, Augmentation Strategy=Multi-Aug, Decoder Configuration=Sem(SegFormer)2023.06 | — | — | — | — | 67.84 | |
| EMSANetBackbone=2x ResNet1012023.06 | — | — | — | — | 66.64 | |
| EMSANetBackbone=2x ResNet34-NBt1D2023.06 | — | — | — | — | 65.57 | |
| HUNetBackbone=HUNet, Evaluation Protocol=Supervised2025.04 | — | 65.5 | — | — | — | |
| Mapillary-trained modelTraining Dataset=Mapillary2021.02 | — | — | — | 1.2 | — | |
| Masked Scene ModelingBackbone=HUNet, Evaluation Protocol=Linear2025.04 | — | 44.4 | — | — | — | |
| MM3DBackbone=PT, Evaluation Protocol=Linear2025.04 | — | 4.3 | — | — | — | |
| MSCBackbone=SparseUNet, Pre-training=Masked Scene Contrast, Downstream Framework=PointGroup2023.03 | — | 59.6 | — | — | — | |
| MSCBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | — | 10.1 | — | — | — | |
| MSCBackbone=HUNet, Evaluation Protocol=Linear2025.04 | — | 24.5 | — | — | — | |
| OESSLBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | — | 13.6 | — | — | — | |
| OpenImages-trained modelTraining Dataset=OpenImages2021.02 | — | — | — | 1.7 | — | |
| PCBackbone=SparseUNet, Pre-training=PointContrast, Downstream Framework=PointGroup2023.03 | — | 58 | — | — | — | |
| SCBackbone=SparseUNet, Pre-training=Train from scratch, Downstream Framework=PointGroup2023.03 | — | 56.9 | — | — | — | |
| ScanNet-trained modelTraining Dataset=ScanNet2021.02 | — | — | — | 35.6 | — | |
| SparseUNet + CSCBackbone=SparseUNet, Pre-training=CSC, Instance Segmentation Framework=PointGroup, Params.=39.2M2023.08 | — | 59.4 | — | — | — | |
| SparseUNet + PCBackbone=SparseUNet, Pre-training=PC, Instance Segmentation Framework=PointGroup, Params.=39.2M2023.08 | — | 58 | — | — | — | |
| SR-UNetBackbone=SR-UNet, Evaluation Protocol=Supervised2025.04 | — | 56.9 | — | — | — | |
| UniDetTraining Dataset=Unified (COCO, CityScapes, Mapillary, VIPER, ScanNet, OpenImages)2021.02 | — | — | — | 28.7 | — | |
| VIPER-trained modelTraining Dataset=VIPER2021.02 | — | — | — | 0 | — |