Loading the SOTA2 catalog…
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding · SOTA2 Research