Loading the SOTA2 catalog…
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model · SOTA2 Research