Loading the SOTA2 catalog…
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models · SOTA2 Research