Loading the SOTA2 catalog…
VCoder: Versatile Vision Encoders for Multimodal Large Language Models · SOTA2 Research