3D Scene Understanding on ScanQA
20.8METEORLLaVA-3D
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaVA-3DInput Modality=3D/2.5D-Input2026.04 | 20.8 | 16.4 | 49.6 | 103.1 | |
| Video-3D-LLMInput Modality=3D/2.5D-Input2026.04 | 20 | 16.4 | 49.3 | 102.1 | |
| VLM-3R-7BInput Modality=Video-Input2026.04 | 19.7 | 15.5 | 49.1 | 101.9 | |
| GeoAlign-4BInput Modality=Video-Input2026.04 | 19.4 | 15.7 | 48.2 | 99.4 | |
| 3D-LLaVAInput Modality=3D/2.5D-Input2026.04 | 18.4 | 17.1 | 43.1 | 92.6 | |
| Spatial-MLLM-4BInput Modality=Video-Input2026.04 | 18.4 | 14.8 | 45 | 91.8 | |
| ChatSceneInput Modality=3D/2.5D-Input2026.04 | 18 | 14.3 | 41.6 | 87.7 | |
| LLaVA-Video-7BInput Modality=Video-Input2026.04 | 17.7 | 3.1 | 44.6 | 88.7 | |
| LL3DAInput Modality=3D/2.5D-Input2026.04 | 15.9 | 13.5 | 37.3 | 76.8 | |
| Oryx-34BInput Modality=Video-Input2026.04 | 15 | — | 37.3 | 72.3 | |
| 3D-LLMInput Modality=3D/2.5D-Input2026.04 | 14.5 | 12 | 35.7 | 69.4 | |
| 3D-VisTAInput Modality=Task-Specific2026.04 | 13.9 | 10.4 | 35.7 | 69.6 | |
| SQA3DInput Modality=Task-Specific2026.04 | 13.5 | 11.2 | 34.5 | — | |
| ScanQAInput Modality=Task-Specific2026.04 | 13.1 | 10.1 | 33.3 | 64.9 | |
| Qwen2.5-VL-72BInput Modality=Video-Input2026.04 | 13 | 12 | 35.2 | 66.9 | |
| Qwen2.5-VL-7BInput Modality=Video-Input2026.04 | 11.4 | 8 | 29.3 | 53.9 |