Loading the SOTA2 catalog…
Leveraging Foundation models for Unsupervised Audio-Visual Segmentation · SOTA2 Research