Scene Spatial Awareness QA on 3D-GRAND
75.32Binary AccuracyGPT-4V
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4VTrainable Params=-, Input=Multi-view Img.2025.12 | 75.32 | 53.68 | |
| Lemon-7BTrainable Params=7.63B, Input=3D Point Cloud2025.12 | 74.32 | 53.45 | |
| GPT-4VTrainable Params=-, Input=Bird-view Img.2025.12 | 71.18 | 53.72 | |
| GPT-4VTrainable Params=-, Input=Single-view Img.2025.12 | 69.23 | 52.34 | |
| Qwen2.5-VL-7BTrainable Params=7.61B, Input=Multi-view Img.2025.12 | 69.1 | 49.3 | |
| LSceneLLMInput=3D Point Cloud2025.12 | 65.46 | 45.79 | |
| Qwen2.5-VL-7BTrainable Params=7.61B, Input=Single-view Img.2025.12 | 64.32 | 47.56 | |
| ShapeLLM-13BTrainable Params=13.04B, Input=3D Point Cloud2025.12 | 60.27 | 42.34 | |
| LLaVA-1.5-13BTrainable Params=13.03B, Input=Multi-view Img.2025.12 | 59.8 | 41.2 | |
| ShapeLLM-7BTrainable Params=7.04B, Input=3D Point Cloud2025.12 | 58.49 | 41.39 | |
| LLaVA-1.5-13BTrainable Params=13.03B, Input=Single-view Img.2025.12 | 57.62 | 40.18 | |
| Ll3daTrainable Params=1.3B, Input=3D Point Cloud2025.12 | 53.45 | 39.6 | |
| 3D-LLMTrainable Params=-, Input=3D Point Cloud2025.12 | 51.25 | 33.43 | |
| LEOTrainable Params=7.01B, Input=3D Point Cloud2025.12 | 49.74 | 30.29 |