Loading the SOTA2 catalog…
VideoBERT: A Joint Model for Video and Language Representation Learning · SOTA2 Research