Loading the SOTA2 catalog…
VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts · SOTA2 Research