Loading the SOTA2 catalog…
LiRA: Learning Visual Speech Representations from Audio through Self-supervision · SOTA2 Research