Loading the SOTA2 catalog…
ViLLa: Video Reasoning Segmentation with Large Language Model · SOTA2 Research