Loading the SOTA2 catalog…
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models · SOTA2 Research