Loading the SOTA2 catalog…
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis · SOTA2 Research