Loading the SOTA2 catalog…
Learning Visual Grounding from Generative Vision and Language Model · SOTA2 Research