3D Semantic Segmentation on ScanNet (val)
79.4mIoUSonata
Evaluation Results
| Method | Links | |||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SonataParam. (M)=124.8, Evaluation Protocol=Full Fine-tuning, Decoder Setting=With decoder2026.04 | 79.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.5 | 86.1 | |
| SonataParam. (M)=108.5, Evaluation Protocol=Full Fine-tuning, Decoder Setting=Without decoder2026.04 | 78.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.3 | 86.1 | |
| PTv3-PPT (sup.)Param. (M)=124.8, Evaluation Protocol=Full Fine-tuning2026.04 | 78.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.3 | 85.9 | |
| Sonata + PointTPAParam. (M)=1.18 (1.09%), Evaluation Protocol=PEFT methods for point cloud, Decoder Setting=Without decoder2026.04 | 78.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.3 | 86.3 | |
| Sonata + DAPTParam. (M)=1.14 (1.06%), Evaluation Protocol=PEFT methods for point cloud, Decoder Setting=Without decoder2026.04 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.9 | 86.5 | |
| Sonata + PointGSTParam. (M)=1.05 (0.97%), Evaluation Protocol=PEFT methods for point cloud, Decoder Setting=Without decoder2026.04 | 77.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.9 | 85.8 | |
| PTv3Param. (M)=124.8, Evaluation Protocol=Training from scratch2026.04 | 77.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92 | 85 | |
| Point-MoE-LParams=100M, Activated=59M, Joint indoor and outdoor training=true2025.05 | 77.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + AdapterParam. (M)=1.90 (1.76%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.7 | 85.1 | |
| PPT-L*Params=98M, Activated=98M, Joint indoor and outdoor training=true, use dataset labels during training=true2025.05 | 76.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OneFormer3D2023.11 | 76.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VMVFInput Modality=Point clouds and images2022.04 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + LoRAParam. (M)=0.89 (0.82%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 76.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.6 | 85.2 | |
| Point-MoE-LParams=100M, Activated=59M, Training Setting=Multi-dataset Joint Training2025.05 | 76 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OctFormervoting=true2023.05 | 75.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Point Transformer V22023.05 | 75.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointTransformerV2Presented at=NeurIPS'222023.11 | 75.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + BitFitParam. (M)=0.12 (0.11%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 75.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.1 | 85.4 | |
| Point-MoE-LParams=100M, Activated=60M, Training Setting=Multi-dataset Joint Training, Precise Evaluator=false2025.05 | 75.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-ScanNetParams=46M, Activated=46M, Training Setting=Single-dataset Training2025.05 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PPT-L*Params=98M, Activated=98M, Training Setting=Multi-dataset Joint Training2025.05 | 74.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PPT-S*Params=47M, Activated=47M, Training Setting=Multi-dataset Joint Training2025.05 | 74.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PPT-LParams=98M, Activated=98M, Training Setting=Multi-dataset Joint Training, Precise Evaluator=false2025.05 | 74.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-LParams=97M, Activated=97M, Joint indoor and outdoor training=true2025.05 | 74.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| O-CNN2023.05 | 74.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OctFormervoting=false2023.05 | 74.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Point-MoE-SParams=59M, Activated=52M, Training Setting=Multi-dataset Joint Training2025.05 | 74.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Point-MoE-SParams=59M, Activated=52M, Training Setting=Multi-dataset Joint Training, Precise Evaluator=false2025.05 | 74.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-ScanNetParams=46M, Activated=46M, Training Setting=Single-dataset Training, Precise Evaluator=false2025.05 | 74.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Stratified Transformer2023.05 | 74.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + RandLoRAParam. (M)=0.71 (0.66%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 74.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.7 | 84.3 | |
| PPT-SParams=47M, Activated=47M, Training Setting=Multi-dataset Joint Training, Precise Evaluator=false2025.05 | 74.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-LParams=97M, Activated=97M, Training Setting=Multi-dataset Joint Training, Mixed Dataset Batch=true, Layernorm=true, Precise Evaluator=false2025.05 | 74 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mix3D2023.05 | 73.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ponder-RGBDBackbone=MinkUNet2023.10 | 73.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + Prefix TuningParam. (M)=0.08 (0.08%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 73.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.3 | 84.2 | |
| Sonata + VeRAParam. (M)=0.08 (0.08%), Evaluation Protocol=General PEFT methods, Decoder Setting=Without decoder2026.04 | 73.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 90.2 | 84 | |
| PTv3-LParams=97M, Activated=97M, Training Setting=Multi-dataset Joint Training, Mixed Dataset Batch=true, Precise Evaluator=false2025.05 | 73.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LargeKernel2023.05 | 73.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CodedVTR (Mink-L)Architecture Type=Transformer, Param=11M2022.03 | 73 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointMetaBase-XXLPresented at=CVPR'232023.11 | 72.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sonata + IDPTParam. (M)=2.61 (2.42%), Evaluation Protocol=PEFT methods for point cloud, Decoder Setting=Without decoder2026.04 | 72.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 89.8 | 82.9 | |
| Minkowski-LArchitecture Type=Convolution, Param=11M2022.03 | 72.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinkowskiNetInput Modality=Colorized point clouds2022.04 | 72.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SparseUNetParam. (M)=39.2, Evaluation Protocol=Training from scratch2026.04 | 72.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 90 | 80.2 | |
| MinkowskiNet42Voxel size=2cm2021.12 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinkowskiNet2023.05 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinkUNetPresented at=CVPR'192023.11 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinkUNet2023.10 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-LParams=97M, Activated=97M, Training Setting=Multi-dataset Joint Training2025.05 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SonataParam. (M)=0.02 (0.02%), Evaluation Protocol=linear probing2026.04 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 89.7 | 83.1 | |
| Fast Point Transformer2023.05 | 72.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OA-CNNsParams=52M, Activated=52M, Training Setting=Multi-dataset Joint Training2025.05 | 71.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointNeXt-XLPresented at=NeurIPS'222023.11 | 71.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointNeXt-XLParam. (M)=41.6, Evaluation Protocol=Training from scratch2026.04 | 71.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvNet (multi-scale head) + CBL (@sub-scenes)CBL @input=false, CBL @sub-scenes=true, multi-scale head=true2022.03 | 71.33 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 89.4 | — | |
| PTv3-SParams=46M, Activated=46M, Training Setting=Multi-dataset Joint Training2025.05 | 71.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| 3D Backbone + DeepViewAggInput Modality=Point clouds and images2022.04 | 71 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvNet + CBL (@sub-scenes)CBL @input=false, CBL @sub-scenes=true, multi-scale head=false2022.03 | 70.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 89.31 | — | |
| Point Transformer2023.05 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointTransformerPresented at=ICCV'212023.11 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PT2023.10 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvNet + CBL (@input)CBL @input=true, CBL @sub-scenes=false, multi-scale head=false2022.03 | 70.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 89.01 | — | |
| PTv3-SParams=46M, Activated=46M, Training Setting=Multi-dataset Joint Training, Mixed Dataset Batch=true, Layernorm=true, Precise Evaluator=false2025.05 | 70 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvNet (multi-scale head)CBL @input=false, CBL @sub-scenes=false, multi-scale head=true2022.03 | 69.83 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 88.88 | — | |
| ConvNetCBL @input=false, CBL @sub-scenes=false, multi-scale head=false2022.03 | 69.71 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 88.97 | — | |
| BPNetInput Modality=Point clouds and images, Supervision=3D supervision only2022.04 | 69.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SparseConvNetVoxel size=2cm2021.12 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPConvInput Modality=Colorized point clouds2022.04 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SparseConvNet2023.05 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SparseConvNet2023.10 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPConv deform2021.12 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| JointPoint2023.05 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPConv2023.05 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPConvPresented at=ICCV'192023.11 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPConv2023.10 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Our 3D BackboneInput Modality=Colorized point clouds2022.04 | 69 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CodedVTR (Mink-M)Architecture Type=Transformer, Param=7M2022.03 | 68.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MVPNetInput Modality=Point clouds and images2022.04 | 68.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PTv3-SParams=46M, Activated=46M, Training Setting=Multi-dataset Joint Training, Mixed Dataset Batch=true, Precise Evaluator=false2025.05 | 67.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointNet++Input Modality=Colorized point clouds2022.04 | 67.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Minkowski-MArchitecture Type=Convolution, Param=7M2022.03 | 67.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VOTR (Mink-L)Architecture Type=Transformer, Param=11M, Note=reproduced2022.03 | 66.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP2SceneFine-tuning Label Percentage=100%2023.01 | 65.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SLidRFine-tuning Label Percentage=100%2023.01 | 64.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointContrastFine-tuning Label Percentage=100%2023.01 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PicassoNetMesh convolution layers (l)=2, Voxel Size / Resolution=4cm2021.03 | 64.3 | 96.1 | 80.4 | 87 | 78.2 | 73.2 | 50.3 | 56.6 | 72 | 59.1 | 90 | 64.1 | 49.9 | 19.2 | 53.9 | 63.9 | 63.7 | 57.1 | 44.9 | 83.5 | 42 | — | — | |
| PPKTFine-tuning Label Percentage=100%2023.01 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointASNL2021.12 | 63.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointASNL2023.05 | 63.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointASNLPresented at=CVPR 202023.11 | 63.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RandomFine-tuning Label Percentage=100%2023.01 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DCM-NetCoordinates / Features=VC,rad/geo2021.03 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PicassoNetMesh convolution layers (l)=2, Voxel Size / Resolution=full2021.03 | 62.5 | 94.4 | 78 | 86.5 | 75.3 | 69.4 | 47.4 | 54.1 | 68.8 | 56.6 | 88.4 | 62.6 | 50.6 | 18.2 | 54.4 | 63.4 | 64.8 | 54.1 | 41.6 | 80.4 | 40.3 | — | — | |
| VOTR (Mink-M)Architecture Type=Transformer, Param=7M, Note=reproduced2022.03 | 62.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SPH3D-GCNVoxel Size / Resolution=3cm2021.03 | 61.2 | 95.5 | 78.3 | 86.9 | 78.9 | 71.4 | 41.6 | 57.1 | 70.6 | 57.4 | 87 | 58.9 | 49.2 | 7.7 | 42.2 | 57.8 | 63.8 | 58.7 | 42.2 | 82.1 | 37.2 | — | — | |
| PointConv2021.12 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointConv2023.05 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointConvSup.=3D, Geometry=true2023.10 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PointConvPresented at=CVPR'192023.11 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |