Loading the SOTA2 catalog…
DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding · SOTA2 Research