3D Object Detection on EmbodiedScan (Vocabulary Breakdown)
0.1907AP (Large-Vocabulary, IoU=0.25)Multi-Modality
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Multi-ModalityInput=RGB-D2023.12 | 0.1907 | 0.5156 | 0.1157 | 0.2815 | 0.2354 | 0.6023 | 0.158 | 0.4499 | 0.0124 | 0.1774 | |
| Depth-OnlyInput=Depth2023.12 | 0.1716 | 0.514 | 0.1052 | 0.2575 | 0.2139 | 0.6114 | 0.1327 | 0.4158 | 0.0274 | 0.2091 | |
| Embodied PerceptronInput=RGB-D2023.12 | 0.1685 | 0.5107 | 0.0977 | 0.2821 | 0.2865 | 0.6751 | 0.1283 | 0.5046 | 0.0709 | 0.3152 | |
| FCAF3D + our decoder + paintingInput=RGB-D2023.12 | 0.151 | 0.5132 | 0.0864 | 0.2666 | 0.2623 | 0.6753 | 0.1139 | 0.5064 | 0.058 | 0.3213 | |
| FCAF3D + our decoderInput=Depth2023.12 | 0.148 | 0.5118 | 0.0877 | 0.2746 | 0.2598 | 0.6712 | 0.1085 | 0.5008 | 0.0572 | 0.3285 | |
| Camera-OnlyInput=RGB2023.12 | 0.128 | 0.3461 | 0.0425 | 0.1307 | 0.174 | 0.4479 | 0.0764 | 0.2422 | 0.0003 | 0.0309 | |
| FCAF3DInput=Depth2023.12 | 0.0907 | 0.4423 | 0.0411 | 0.2022 | 0.1654 | 0.6138 | 0.0673 | 0.4277 | 0.0267 | 0.2483 | |
| ImVoxelNetInput=RGB2023.12 | 0.0615 | 0.2039 | 0.0241 | 0.0631 | 0.1096 | 0.3429 | 0.0412 | 0.154 | 0.0263 | 0.0921 | |
| VoteNetInput=Depth2023.12 | 0.032 | 0.0611 | 0.0038 | 0.0122 | 0.0631 | 0.1226 | 0.0181 | 0.0334 | 0.01 | 0.0183 |