Loading the SOTA2 catalog…
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation · SOTA2 Research