3D Scene Understanding on SQA3D
60.7EM-1VLM-3R-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| VLM-3R-7BInput Modality=Video-Input2026.04 | 60.7 | |
| GeoAlign-4BInput Modality=Video-Input2026.04 | 60.3 | |
| LLaVA-3DInput Modality=3D/2.5D-Input2026.04 | 60.1 | |
| Video-3D-LLMInput Modality=3D/2.5D-Input2026.04 | 58.6 | |
| Spatial-MLLM-4BInput Modality=Video-Input2026.04 | 55.9 | |
| ChatSceneInput Modality=3D/2.5D-Input2026.04 | 54.6 | |
| 3D-LLaVAInput Modality=3D/2.5D-Input2026.04 | 54.5 | |
| Oryx-34BInput Modality=Video-Input2026.04 | 50.9 | |
| 3D-VisTAInput Modality=Task-Specific2026.04 | 48.5 | |
| LLaVA-Video-7BInput Modality=Video-Input2026.04 | 48.5 | |
| ScanQAInput Modality=Task-Specific2026.04 | 47.2 | |
| Qwen2.5-VL-72BInput Modality=Video-Input2026.04 | 47 | |
| SQA3DInput Modality=Task-Specific2026.04 | 46.6 | |
| Qwen2.5-VL-7BInput Modality=Video-Input2026.04 | 46.5 |