Loading the SOTA2 catalog…
Visual Representation Alignment for Multimodal Large Language Models · SOTA2 Research