Referring 3D Instance Segmentation on ScanRefer (val)
74.6mIoUReason3D
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Reason3DInput Modality=ground-truth point clouds2026.01 | 74.6 | 88.4 | 84.2 | |
| 3D-STMNInput Modality=ground-truth point clouds2026.01 | 74.5 | 89.3 | 84 | |
| RG-SANInput Modality=ground-truth point clouds2026.01 | 74.5 | 89.2 | 84.3 | |
| MVGGTInput Modality=sparse RGB inputs2026.01 | 65.2 | 83.6 | 74.5 | |
| TGNNInput Modality=ground-truth point clouds2026.01 | 50.7 | 69.3 | 57.8 | |
| PAR3Dcategory=Generalist 3D-MLLMs2026.06 | 49.9 | — | — | |
| RG-SANInput Modality=ground-truth point clouds2026.01 | 44.6 | 61.7 | 44.9 | |
| MoE3DModality=PC, Model Category=3D LMM2025.11 | 44.4 | — | — | |
| 3D-LLaVAModality=PC, Model Type=3D LMM2025.01 | 43.3 | — | — | |
| 3D-LLaVAModality=PC, Model Category=3D LMM2025.11 | 43.3 | — | — | |
| 3D-LLaVAInput Modality=ground-truth point clouds2026.01 | 43.3 | — | — | |
| 3D-LLaVAcategory=Generalist 3D-MLLMs2026.06 | 43.3 | — | — | |
| Reason3DInput Modality=ground-truth point clouds2026.01 | 42 | 57.9 | 41.9 | |
| Reason3Dcategory=Generalist 3D-MLLMs2026.06 | 42 | — | — | |
| SegPointModality=PC, Model Type=Finetuned 3D LMM2025.01 | 41.7 | — | — | |
| SegPointModality=PC, Model Category=Finetuned 3D LMM2025.11 | 41.7 | — | — | |
| SegPointInput Modality=ground-truth point clouds2026.01 | 41.7 | — | — | |
| SegPointcategory=Finetuned 3D-MLLMs2026.06 | 41.7 | — | — | |
| MVGGTInput Modality=sparse RGB inputs2026.01 | 39.9 | 55.9 | 41.5 | |
| 3D-STMNModality=PC, Model Type=Specialist Model2025.01 | 39.5 | — | — | |
| 3D-STMNModality=PC, Model Category=Specialist Model2025.11 | 39.5 | — | — | |
| 3D-STMNInput Modality=ground-truth point clouds2026.01 | 39.5 | 54.6 | 39.8 | |
| 3D-STMNcategory=Specialist Models2026.06 | 39.5 | — | — | |
| M3DRef-CLIPModality=PC, Model Type=Specialist Model2025.01 | 35.7 | — | — | |
| M3DRef-CLIPModality=PC, Model Category=Specialist Model2025.11 | 35.7 | — | — | |
| M3DRef-CLIPcategory=Specialist Models2026.06 | 35.7 | — | — | |
| LESSStage=Single Stage, Backbone=RoBERTa, Label Effort=< 2 min, Supervision=Mask2024.10 | 33.74 | 53.23 | 29.88 | |
| LESSInput Modality=ground-truth point clouds2026.01 | 33.7 | 53.2 | 29.9 | |
| LESSStage=Single Stage, Backbone=BERT, Label Effort=< 2 min, Supervision=Mask2024.10 | 32.44 | 51.41 | 29.02 | |
| LESSStage=Single Stage, Backbone=GRU, Label Effort=< 2 min, Supervision=Mask2024.10 | 32.19 | 51 | 26.41 | |
| 2D-LiftInput Modality=sparse RGB inputs2026.01 | 31.5 | 51.6 | 26 | |
| X-RefSegStage=Two Stage, Backbone=BERT, Label Effort=> 20 min, Supervision=Ins.+ Sem.2024.10 | 29.94 | 40.33 | 33.77 | |
| X-RefSeg3DModality=PC, Model Type=Specialist Model2025.01 | 29.9 | — | — | |
| X-RefSeg3DModality=PC, Model Category=Specialist Model2025.11 | 29.9 | — | — | |
| X-RefSegStage=Two Stage, Backbone=GRU, Label Effort=> 20 min, Supervision=Ins.+ Sem.2024.10 | 29.77 | 39.85 | 33.52 | |
| TGNNInput Modality=ground-truth point clouds2026.01 | 28.8 | 38.6 | 32.7 | |
| TGNNStage=Two Stage, Backbone=BERT, Label Effort=> 20 min, Supervision=Ins.+ Sem.2024.10 | 27.8 | 37.5 | 31.4 | |
| TGNNModality=PC, Model Type=Specialist Model2025.01 | 27.8 | — | — | |
| TGNNModality=PC, Model Category=Specialist Model2025.11 | 27.8 | — | — | |
| two-stageInput Modality=sparse RGB inputs2026.01 | 27.5 | 43.8 | 20.9 | |
| TGNNStage=Two Stage, Backbone=GRU, Label Effort=> 20 min, Supervision=Ins.+ Sem.2024.10 | 26.1 | 35 | 29 | |
| two-stageInput Modality=sparse RGB inputs2026.01 | 18.5 | 28.7 | 13.9 | |
| 2D-LiftInput Modality=sparse RGB inputs2026.01 | 17.8 | 27.3 | 9.6 |