3D Visual Grounding on Multi3DRefer
61.6Accuracy@0.25Vid-LLM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Vid-LLMModel Category=Video-based models2025.09 | 61.6 | 56.1 | |
| 3DRSModel Category=3D-based models2025.09 | 60.4 | 54.9 | |
| Video-3D LLMModel Category=3D-based models2025.09 | 58 | 52.7 | |
| ChatSceneModel Category=3D-based models2025.09 | 57.1 | 52.4 | |
| 3DRSModel Category=Video-based models, 3D Geometry Source=VGGT-generated2025.09 | 52.8 | 48.3 | |
| LLaVA-3DModel Category=3D-based models2025.09 | 49.8 | 43.6 | |
| Grounded3D-LLMModel Category=3D-based models2025.09 | 44.7 | 40.8 |