Loading the SOTA2 catalog…
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation · SOTA2 Research