Loading the SOTA2 catalog…
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction · SOTA2 Research