Loading the SOTA2 catalog…
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models · SOTA2 Research