Loading the SOTA2 catalog…
VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs · SOTA2 Research