Loading the SOTA2 catalog…
A vector quantized masked autoencoder for audiovisual speech emotion recognition · SOTA2 Research