Loading the SOTA2 catalog…
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models · SOTA2 Research