Open-vocabulary Semantic Segmentation on nuScenes (val)
51.1mIoUIGLOSSmix
Evaluation Results
| Method | Links | |
|---|---|---|
| IGLOSSmixOpen Vocabulary Status=fully OV, Input Modality=3D points only, VFM Ensembling=Yes2026.04 | 51.1 | |
| IGLOSSOpen Vocabulary Status=fully OV, Input Modality=3D points only, VFM Ensembling=No2026.04 | 47.5 | |
| SASOpen Vocabulary Status=OV with vocabulary/prompt bias at training, Input Modality=3D points only, VFM Ensembling=Yes2026.04 | 47.5 | |
| GGSDOpen Vocabulary Status=OV with vocabulary/prompt bias at training, Input Modality=3D points only, VFM Ensembling=No2026.04 | 46.1 | |
| OV3D with OpenSceneOpen Vocabulary Status=OV with captioner bias at training, Input Modality=3D points only, VFM Ensembling=Yes2026.04 | 45.5 | |
| OV3DOpen Vocabulary Status=OV with captioner bias at training, Input Modality=3D points only, VFM Ensembling=No2026.04 | 44.6 | |
| OpenScene (I+L)Open Vocabulary Status=fully OV, Input Modality=Includes images, VFM Ensembling=No2026.04 | 42.1 | |
| OpenScene (L)Open Vocabulary Status=fully OV, Input Modality=3D points only, VFM Ensembling=No2026.04 | 41.3 | |
| 3D-AVS (I+L)Open Vocabulary Status=OV with captioner bias at training, Input Modality=Includes images, VFM Ensembling=No2026.04 | 36.2 | |
| SALOpen Vocabulary Status=fully OV, Input Modality=3D points only, VFM Ensembling=No2026.04 | 33.9 | |
| 3D-AVS (L)Open Vocabulary Status=OV with captioner bias at training, Input Modality=3D points only, VFM Ensembling=No2026.04 | 33.4 | |
| CLIP2SceneOpen Vocabulary Status=OV with vocabulary/prompt bias at training, Input Modality=3D points only, VFM Ensembling=No2026.04 | 20.8 | |
| MaskCLIP→3DOpen Vocabulary Status=fully OV, Input Modality=Includes images, VFM Ensembling=No2026.04 | 16.6 |