Loading the SOTA2 catalog…
Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models · SOTA2 Research