Loading the SOTA2 catalog…
VISA: Reasoning Video Object Segmentation via Large Language Models · SOTA2 Research