Loading the SOTA2 catalog…
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation · SOTA2 Research