Loading the SOTA2 catalog…
FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks · SOTA2 Research