Loading the SOTA2 catalog…
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation · SOTA2 Research