Loading the SOTA2 catalog…
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues · SOTA2 Research