Loading the SOTA2 catalog…
Grounding Language Models to Images for Multimodal Inputs and Outputs · SOTA2 Research