Loading the SOTA2 catalog…
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention · SOTA2 Research