Loading the SOTA2 catalog…
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation · SOTA2 Research