Loading the SOTA2 catalog…
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues · SOTA2 Research