Loading the SOTA2 catalog…
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models · SOTA2 Research