Loading the SOTA2 catalog…
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning · SOTA2 Research