Loading the SOTA2 catalog…
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations · SOTA2 Research