Loading the SOTA2 catalog…
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation · SOTA2 Research