Loading the SOTA2 catalog…
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model · SOTA2 Research