Loading the SOTA2 catalog…
Small Vision-Language Models are Smart Compressors for Long Video Understanding · SOTA2 Research