Loading the SOTA2 catalog…
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks · SOTA2 Research