Loading the SOTA2 catalog…
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts · SOTA2 Research