Loading the SOTA2 catalog…
Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants · SOTA2 Research