Loading the SOTA2 catalog…
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models · SOTA2 Research