3D Dense Captioning on Scan2Cap (test)
86.1CIDEr@0.53DRS
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| 3DRSModel Type=3D-based models2025.09 | 86.1 | — | — | — | — | 41.6 | 28.9 | 62.3 | |
| LLaVA-3DModel Type=3D-based models2025.09 | 84.1 | — | — | — | — | 42.6 | 29 | 63.4 | |
| Video-3D LLMModel Type=3D-based models2025.09 | 83.8 | — | — | — | — | 42.4 | 28.9 | 62.3 | |
| Vid-LLMModel Type=Video-based models2025.09 | 81.5 | — | — | — | — | 40.9 | 28.7 | 61.8 | |
| VGLLMModel Type=Video-based models2025.09 | 78.6 | — | — | — | — | 40.9 | 28.6 | 62.4 | |
| ChatSceneModel Type=3D-based models2025.09 | 77.1 | — | — | — | — | 36.3 | 28 | 58.1 | |
| 3DRSModel Type=Video-based models, VGGT-generated 3D geometry=true2025.09 | 76.4 | — | — | — | — | 39.6 | 27.3 | 57.1 | |
| Grounded3D-LLMModel Type=3D-based models2025.09 | 70.2 | — | — | — | — | 35 | — | — | |
| LEOModel Type=3D-based models2025.09 | 68.4 | — | — | — | — | 36.9 | 27.7 | 57.8 | |
| Scan2CapModel Type=3D-based models2025.09 | 39.1 | — | — | — | — | 23.3 | 22 | 44.8 | |
| 3D-VisTAIoU threshold=0.5, Training Protocol=single dataset2024.05 | — | 66.9 | 34 | 27.1 | 54.3 | — | — | — | |
| 3DJCGIoU threshold=0.5, Training Protocol=single dataset2024.05 | — | 47.7 | 31.5 | 24.3 | 51.8 | — | — | — | |
| PQ3DIoU threshold=0.5, Training Protocol=single dataset2024.05 | — | 75.6 | 34.4 | 28.6 | 57.1 | — | — | — | |
| PQ3DIoU threshold=0.5, Training Protocol=unified training2024.05 | — | 80.3 | 36 | 29.1 | 57.9 | — | — | — | |
| Scan2CapIoU threshold=0.5, Training Protocol=single dataset2024.05 | — | 35.2 | 22.4 | 21.4 | 43.5 | — | — | — |