Loading the SOTA2 catalog…
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning · SOTA2 Research