Loading the SOTA2 catalog…
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception · SOTA2 Research