Loading the SOTA2 catalog…
VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks · SOTA2 Research