Loading the SOTA2 catalog…
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation · SOTA2 Research