Loading the SOTA2 catalog…
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding · SOTA2 Research