Loading the SOTA2 catalog…
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks · SOTA2 Research