3D referring expression grounding on GAPartNet (test)
7.6Drawer AccuracySHAPELLM-13B
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| SHAPELLM-13BInput=3D Point Cloud2024.02 | 7.6 | 26.7 | 11.5 | 6.7 | 6.8 | 11.1 | 11.7 | |
| SHAPELLM-7BInput=3D Point Cloud2024.02 | 5.9 | 25.8 | 11.5 | 3.4 | 5.1 | 11.1 | 10.5 | |
| LLaVA-13BInput=4-View 2D Image, Fine-tuned on GAPartNet images=true2024.02 | 2.5 | 13.7 | 7.7 | 0 | 4.3 | 11.1 | 6.2 | |
| LLaVA-13BInput=1-View 2D Image, Fine-tuned on GAPartNet images=true2024.02 | 1.8 | 9.3 | 3.8 | 0 | 2.1 | 11.1 | 4.4 | |
| GPT-4VInput=4-View 2D Image, In-context demonstrations=32024.02 | 0.1 | 1.6 | 0 | 0 | 0 | 0 | 0.3 | |
| LLaVA-13BInput=1-View 2D Image, Fine-tuned on GAPartNet images=false2024.02 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| LLaVA-13BInput=4-View 2D Image, Fine-tuned on GAPartNet images=false2024.02 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| GPT-4VInput=4-View 2D Image, In-context demonstrations=02024.02 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |