Loading the SOTA2 catalog…
End-to-End Multimodal Representation Learning for Video Dialog · SOTA2 Research