Loading the SOTA2 catalog…
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality · SOTA2 Research