3D Visual Grounding on Multi3DRefer (val)
59.8F1@0.50GPT4Scene (HDM)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT4Scene (HDM)Vision modal.=B,I, # of tokens=80002026.05 | 59.8 | 64.5 | — | — | |
| Proxy3DVision modal.=I, # of tokens=700, Backbone=Qwen2.5-VL-7B2026.05 | 57.5 | 62 | — | — | |
| APEIRIAOutput=Text, modular inference enhancement=true2026.05 | 55.2 | 60.9 | — | — | |
| Descrip3DVision modal.=I2026.05 | 55.1 | 59.4 | — | — | |
| 3DRSVision modal.=I, # of tokens=80002026.05 | 54.9 | 60.4 | — | — | |
| CVPMethod Category=3D LMMs2025.12 | 54.7 | 60.2 | — | — | |
| APEIRIAOutput=Text, modular inference enhancement=false2026.05 | 53.8 | 59.2 | — | — | |
| Inst3D-LMMMethod Paradigm=Generalist2025.03 | 53.5 | 58.3 | — | — | |
| Inst3D-LMMOutput=Text, modular inference enhancement=false2026.05 | 53.5 | 58.3 | — | — | |
| Video-3D-LLMMethod Category=3D LMMs2025.12 | 52.7 | 58 | — | — | |
| Video-3D-LLMVision modal.=I, # of tokens=80002026.05 | 52.7 | 58 | — | — | |
| Video-3D LLMOutput=Head, modular inference enhancement=false2026.05 | 52.7 | 58 | — | — | |
| Chat-SceneMethod Paradigm=Generalist2025.03 | 52.4 | 57.1 | — | — | |
| Chat-SceneMethod Category=3D LMMs2025.12 | 52.4 | 57.1 | — | — | |
| Chat-SceneVision modal.=P2026.05 | 52.4 | 57.1 | — | — | |
| Chat-SceneOutput=Text, modular inference enhancement=false2026.05 | 52.4 | 57.1 | — | — | |
| PQ3DOutput=Head, modular inference enhancement=false2026.05 | 50.1 | — | — | — | |
| LLaVA-3DMethod Category=3D LMMs2025.12 | 43.6 | 49.8 | — | — | |
| LLaVA-3DOutput=Head, modular inference enhancement=false2026.05 | 43.6 | 49.8 | — | — | |
| Chat-3D v2Method Paradigm=Specialist2025.03 | 41.6 | 45.1 | — | — | |
| Grounded 3D-LLMOutput=Head, modular inference enhancement=false2026.05 | 40.8 | 44.7 | — | — | |
| Grounded 3D-LLMMethod Paradigm=Generalist2025.03 | 40.6 | 45.2 | — | — | |
| Grounded 3D-LLMMethod Category=3D LMMs2025.12 | 40.6 | 45.2 | — | — | |
| M3DRef-CLIPMethod Paradigm=Closed-set, full sup.2025.03 | 38.4 | 42.8 | — | — | |
| M3DRef-CLIPMethod Category=Expert models2025.12 | 38.4 | 42.8 | — | — | |
| 3DJCGMethod Category=Expert models2025.12 | 26.6 | — | — | — | |
| 3DVG-TransMethod Paradigm=Closed-set, full sup.2025.03 | 25.5 | 30.2 | — | — | |
| 3DVG-TransMethod Category=Expert models2025.12 | 25.5 | — | — | — | |
| Qwen2-VL-7BVision modal.=I, # of tokens=80002026.05 | 19.9 | 21.1 | — | — | |
| 3D-LLMModel Type=Single-Stage 3D LMMs2024.09 | — | — | 42.7 | — | |
| Chat-SceneModel Type=Two-Stage 3D LMMs2024.09 | — | — | 57.1 | 52.4 | |
| Grounded 3D-LLMModel Type=Two-Stage 3D LMMs2024.09 | — | — | 45.2 | 40.6 | |
| LLaVA-3DModel Type=Two-Stage 3D LMMs, Protocol=Two-stage application2024.09 | — | — | 68.1 | 62.9 | |
| LLaVA-3DModel Type=Single-Stage 3D LMMs2024.09 | — | — | 49.8 | 43.6 | |
| M3DRef-CLIPModel Type=Task-specific2024.09 | — | — | 42.8 | 38.4 | |
| M3DRef-CLIPOutput=Head, modular inference enhancement=false2026.05 | — | 42.8 | — | — |