Instance Segmentation on ScanNet200 (val)
45.3mAP@50ODIN-Swin-B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ODIN-Swin-BInput Modality=Sensor RGBD Point Cloud, Backbone=Swin-B2024.01 | 45.3 | 53.1 | 31.5 | — | — | — | |
| OneFormer3D2023.11 | 40.8 | 45.4 | 30.6 | — | — | — | |
| PQ3DVocabulary setting=Closed-vocabulary, Classification head=closed-vocabulary2024.05 | 38.9 | 46.3 | 27 | 35.8 | 24.2 | 20 | |
| MAFTInput Modality=Mesh Sampled Point Cloud2024.01 | 38.2 | 43.3 | 29.2 | — | — | — | |
| QueryFormerInput Modality=Mesh Sampled Point Cloud2024.01 | 37.1 | 43.4 | 28.1 | — | — | — | |
| Mask3D2023.11 | 37 | 42.3 | 27.4 | — | — | — | |
| Mask3DInput Modality=Mesh Sampled Point Cloud2024.01 | 37 | 42.3 | 27.4 | — | — | — | |
| ODIN-ResNet50Input Modality=Sensor RGBD Point Cloud, Backbone=ResNet502024.01 | 36.9 | 43.8 | 25.6 | — | — | — | |
| Mask3DVocabulary setting=Closed-vocabulary2024.05 | 36.2 | 41.4 | 26.9 | 39.8 | 21.7 | 17.9 | |
| Sonata (full)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 35.6 | 42.1 | 25.4 | — | — | — | |
| TD3D2023.11 | 34.8 | 40.4 | 23.1 | — | — | — | |
| PointINSProtocol=Fine-tuning2026.03 | 34.8 | 41.5 | 25.1 | — | — | — | |
| PPT (sup.)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 34.1 | 40.8 | 24 | — | — | — | |
| PTv3 (sup.)Protocol=Supervised2026.03 | 34.1 | 40.8 | 24 | — | — | — | |
| MSC (full)Params=124.8M, Learn.=124.8M, Pct.=100%, Evaluation Protocol=full fine-tuning2025.03 | 33.8 | 40.5 | 23.4 | — | — | — | |
| ChorusEvaluation Protocol=fine-tuning2025.12 | 33.7 | 39.3 | — | — | — | — | |
| DOSProtocol=Fine-tuning2026.03 | 33.5 | 38.6 | 22.5 | — | — | — | |
| Sonata (dec.)Params=124.8M, Learn.=16.3M, Pct.=13%, Evaluation Protocol=decoder probing2025.03 | 33.3 | 40.8 | 22.8 | — | — | — | |
| PointGroupParams=124.8M, Learn.=124.8M, Pct.=100%2025.03 | 33.2 | 40.1 | 23.1 | — | — | — | |
| HUNetBackbone=HUNet, Evaluation Protocol=Supervised2025.04 | 32.8 | — | — | — | — | — | |
| PTv3Evaluation Protocol=supervised2025.12 | 32.3 | 40.1 | — | — | — | — | |
| PointINSProtocol=Decoder2026.03 | 32.2 | 38.3 | 21.3 | — | — | — | |
| ChorusEvaluation Protocol=decoder probing2025.12 | 31.8 | 38.8 | — | — | — | — | |
| SonataEvaluation Protocol=fine-tuning2025.12 | 31.5 | 38.3 | — | — | — | — | |
| PonderV2Baseline=PointGroup, Pre-training=PonderV22025.06 | 30.5 | 37.6 | 20.1 | — | — | — | |
| PPT (f.t.)Baseline=PointGroup, Pre-training=PPT, Protocol=fine-tuned2025.06 | 29.4 | 36.8 | 19.4 | — | — | — | |
| SonataEvaluation Protocol=decoder probing2025.12 | 29.3 | 36.2 | — | — | — | — | |
| UniPre3DBaseline=PointGroup, Pre-training=UniPre3D2025.06 | 29.2 | 37.1 | 18.7 | — | — | — | |
| SegContrastProtocol=Fine-tuning2026.03 | 29.2 | 36.5 | 19.6 | — | — | — | |
| DOSProtocol=Decoder2026.03 | 28.5 | 36.2 | 18.8 | — | — | — | |
| SonataProtocol=Fine-tuning2026.03 | 28.5 | 35.8 | 20 | — | — | — | |
| NOMAEProtocol=Fine-tuning2026.03 | 28.4 | 36 | 19.3 | — | — | — | |
| PonderV2+Baseline=PointGroup, Pre-training=PonderV2+2025.06 | 28.3 | 36 | 18.4 | — | — | — | |
| PQ3DVocabulary setting=Open-vocabulary, Mode=promptable2024.05 | 28 | 32.5 | 20.2 | 30.9 | 17 | 11.3 | |
| PSAProtocol=Fine-tuning2026.03 | 28 | 35.6 | 19.4 | — | — | — | |
| SonataProtocol=Decoder2026.03 | 27.8 | 34.8 | 17.9 | — | — | — | |
| MSCBaseline=PointGroup, Pre-training=MSC2025.06 | 26.8 | 34.3 | 17.3 | — | — | — | |
| MSCBackbone=SparseUNet, Pre-training=Masked Scene Contrast, Downstream Framework=PointGroup2023.03 | 26.8 | — | — | — | — | — | |
| LGroundBaseline=PointGroup, Pre-training=LGround2025.06 | 26.1 | — | — | — | — | — | |
| PointGroup + LGroundcontext_enhancement=LGround2023.11 | 26.1 | — | — | — | — | — | |
| CSCBaseline=PointGroup, Pre-training=CSC2025.06 | 25.2 | — | — | — | — | — | |
| CSCBackbone=SparseUNet, Pre-training=Contrastive Scene Contexts, Downstream Framework=PointGroup2023.03 | 25.2 | — | — | — | — | — | |
| PCBaseline=PointGroup, Pre-training=PC2025.06 | 24.9 | — | — | — | — | — | |
| PCBackbone=SparseUNet, Pre-training=PointContrast, Downstream Framework=PointGroup2023.03 | 24.9 | — | — | — | — | — | |
| PointINSProtocol=Linear Probing2026.03 | 24.9 | 35.8 | 13.4 | — | — | — | |
| NOMAEProtocol=Decoder2026.03 | 24.8 | 31.8 | 15.8 | — | — | — | |
| NoneBaseline=PointGroup, Pre-training=None2025.06 | 24.5 | 32.2 | 15.8 | — | — | — | |
| SCBackbone=SparseUNet, Pre-training=Train from scratch, Downstream Framework=PointGroup2023.03 | 24.5 | — | — | — | — | — | |
| PointGroup2023.11 | 24.5 | — | — | — | — | — | |
| SR-UNetBackbone=SR-UNet, Evaluation Protocol=Supervised2025.04 | 24.5 | — | — | — | — | — | |
| ChorusEvaluation Protocol=linear probing2025.12 | 21.9 | 31.6 | — | — | — | — | |
| Mask3DInput Modality=Sensor RGBD Point Cloud, trained_by_authors=true2024.01 | 21.4 | 24.3 | 15.5 | — | — | — | |
| Sonata (lin.)Params=124.8M, Learn.=<0.2M, Pct.=<0.2%, Evaluation Protocol=linear probing2025.03 | 21.3 | 30.9 | 10.9 | — | — | — | |
| SonataEvaluation Protocol=linear probing2025.12 | 20.9 | 30 | — | — | — | — | |
| DOSProtocol=Linear Probing2026.03 | 20.6 | 27.9 | 10.9 | — | — | — | |
| OpenMask3DInput Modality=Zero-Shot2024.01 | 19.9 | 23.1 | 15.4 | — | — | — | |
| OpenMask3DVocabulary setting=Open-vocabulary2024.05 | 19.9 | 23.1 | 15.4 | 17.1 | 14.1 | 14.9 | |
| SegContrastProtocol=Decoder2026.03 | 18 | 25.5 | 10.5 | — | — | — | |
| PSAProtocol=Decoder2026.03 | 18 | 25.5 | 10.9 | — | — | — | |
| SonataProtocol=Linear Probing2026.03 | 17.5 | 25.2 | 8.7 | — | — | — | |
| OpenSceneVocabulary setting=Open-vocabulary2024.05 | 15.2 | 17.8 | 11.7 | 13.4 | 11.6 | 9.9 | |
| Masked Scene ModelingBackbone=HUNet, Evaluation Protocol=Linear2025.04 | 8.8 | — | — | — | — | — | |
| NOMAEProtocol=Linear Probing2026.03 | 6.7 | 12.7 | 2.9 | — | — | — | |
| PSAProtocol=Linear Probing2026.03 | 5.3 | 11.4 | 2.4 | — | — | — | |
| SegContrastProtocol=Linear Probing2026.03 | 4.7 | 10.2 | 1.8 | — | — | — | |
| MSCBackbone=HUNet, Evaluation Protocol=Linear2025.04 | 3.8 | — | — | — | — | — | |
| OESSLBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | 2.5 | — | — | — | — | — | |
| MSCBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | 1.6 | — | — | — | — | — | |
| MSCEvaluation Protocol=linear probing2025.12 | 1 | 2.3 | — | — | — | — | |
| MSC (lin.)Params=124.8M, Learn.=<0.2M, Pct.=<0.2%, Evaluation Protocol=linear probing2025.03 | 1 | 2.3 | 0.4 | — | — | — | |
| MM3DBackbone=PT, Evaluation Protocol=Linear2025.04 | 0.4 | — | — | — | — | — | |
| CSCBackbone=SR-UNet, Evaluation Protocol=Linear2025.04 | 0.1 | — | — | — | — | — | |
| LAM3CVGPC Pretraining Data=16k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | — | — | 5.8 | — | — | — | |
| LAM3CVGPC Pretraining Data=16k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 18.7 | — | — | — | |
| LAM3CVGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | — | — | 8.3 | — | — | — | |
| LAM3CVGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 19.6 | — | — | — | |
| LAM3C*VGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | — | — | 9.5 | — | — | — | |
| LAM3C*VGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | — | — | 21.9 | — | — | — | |
| LAM3C*Real Pretraining Data=15k, VGPC Pretraining Data=49k, Evaluation Protocol=LP, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | — | — | 10.8 | — | — | — | |
| LAM3C*Real Pretraining Data=15k, VGPC Pretraining Data=49k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Large), Pre-training steps=434k2025.12 | — | — | 21.3 | — | — | — | |
| MSCReal Pretraining Data=7k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 23.4 | — | — | — | |
| PPTReal Pretraining Data=1k, Synth Pretraining Data=21k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 24 | — | — | — | |
| PTv3Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | — | — | 0.02 | — | — | — | |
| PTv3Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 17.2 | — | — | — | |
| SonataReal Pretraining Data=18k, Synth Pretraining Data=121k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 25.4 | — | — | — | |
| Sonata (all real)Real Pretraining Data=15k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | — | — | 9.7 | — | — | — | |
| Sonata (all real)Real Pretraining Data=15k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 20.8 | — | — | — | |
| Sonata (ScanNet)Real Pretraining Data=1k, Evaluation Protocol=LP, Backbone=PTv3 (Base)2025.12 | — | — | 8.6 | — | — | — | |
| Sonata (ScanNet)Real Pretraining Data=1k, Evaluation Protocol=Full-FT, Backbone=PTv3 (Base)2025.12 | — | — | 19.7 | — | — | — |