Loading the SOTA2 catalog…
COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval · SOTA2 Research