Loading the SOTA2 catalog…
HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval · SOTA2 Research