Loading the SOTA2 catalog…
VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding · SOTA2 Research