Loading the SOTA2 catalog…
When LLaVA Meets Objects: Token Composition for Vision-Language-Models · SOTA2 Research