Loading the SOTA2 catalog…
$\beta$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment · SOTA2 Research