3D Semantic Segmentation on ScanNet (mIoU, mAcc)
72mIoUFully supervised
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Fully supervisedTest-Time Images=✗, Training Paradigm=Fully supervised upper bound2026.06 | 72 | 80.7 | |
| Casper3D + 3D DistillationTest-Time Images=✗, Training Paradigm=3D point-cloud training2026.06 | 65.8 | 76.7 | |
| Casper3DTest-Time Images=✓, Training Paradigm=2D fusion w/ training2026.06 | 64.6 | 75.5 | |
| PLOVIS2026.04 | 60.1 | 73.7 | |
| PGOV3DTest-Time Images=✗, Training Paradigm=3D point-cloud training2026.06 | 59.5 | 73.2 | |
| OV3DTest-Time Images=✗, Training Paradigm=3D point-cloud training2026.06 | 57.3 | 72.9 | |
| MLP probingFine-tuning strategy=MLP probing2026.04 | 52.9 | 64.9 | |
| DG-Net2026.04 | 52.9 | 67.2 | |
| OpenScene + 3D DistillationTest-Time Images=✗, Training Paradigm=3D point-cloud training2026.06 | 52.9 | 63.2 | |
| AAD-Net2026.04 | 52.8 | 65.1 | |
| ERDA2026.04 | 52.8 | 66.3 | |
| OpenScene-2DTest-Time Images=✓, Training Paradigm=2D fusion Zero-shot2026.06 | 50 | 62.7 | |
| Decoder probingFine-tuning strategy=Decoder probing2026.04 | 49.9 | 62.5 | |
| Full fine-tuningFine-tuning strategy=Full fine-tuning2026.04 | 47 | 61.1 | |
| MSeg VotingTest-Time Images=✓, Training Paradigm=2D fusion Zero-shot2026.06 | 45.6 | 54.4 | |
| CLIP-FO3DTest-Time Images=✗, Training Paradigm=3D point-cloud training2026.06 | 30.2 | 49.1 | |
| CLIP-FO3D, feature projectionTest-Time Images=✓, Training Paradigm=2D fusion Zero-shot2026.06 | 27.6 | 47.7 | |
| MaskCLIP-3DTest-Time Images=✓, Training Paradigm=2D fusion Zero-shot2026.06 | 9.7 | 21.6 |