Loading the SOTA2 catalog…
Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling · SOTA2 Research