3D Multimodal Comprehension on 3D MM-Vet (test)
65.1Recognition AccuracyGPT-4V
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-4VInput=4-View 2D Image, Zero-shot=true2024.02 | 65.1 | 69.1 | 61.4 | 52.9 | 65.5 | 63.4 | |
| GPT-4VInput=1-View 2D Image, Zero-shot=true2024.02 | 53.7 | 59.5 | 61.1 | 54.7 | 59 | 57.4 | |
| SHAPELLM-13BInput=3D Point Cloud, Zero-shot=true2024.02 | 46.8 | 53 | 53.9 | 45.3 | 68.4 | 53.1 | |
| PointLLM-13BInput=3D Point Cloud, Zero-shot=true2024.02 | 46.6 | 48.3 | 38.8 | 45.2 | 50.9 | 46.6 | |
| SHAPELLM-7BInput=3D Point Cloud, Zero-shot=true2024.02 | 45.7 | 42.7 | 43.4 | 39.9 | 64.5 | 47.4 | |
| DreamLLM-7BInput=4-View 2D Image, Zero-shot=true2024.02 | 42.2 | 54.4 | 50.8 | 48.9 | 54.5 | 50.3 | |
| PointLLM-7BInput=3D Point Cloud, Zero-shot=true2024.02 | 40.6 | 49.5 | 34.3 | 29.1 | 48.7 | 41.2 | |
| LLAVA-13BInput=1-View 2D Image, Zero-shot=true2024.02 | 40 | 55.3 | 51.3 | 43.2 | 51.1 | 47.9 | |
| PointBind&LLMInput=3D Point Cloud, Zero-shot=true2024.02 | 16.9 | 13 | 18.5 | 32.9 | 40.4 | 23.5 |