Loading the SOTA2 catalog…
Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping · SOTA2 Research