Loading the SOTA2 catalog…
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks · SOTA2 Research