Loading the SOTA2 catalog…
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition · SOTA2 Research