Loading the SOTA2 catalog…
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment · SOTA2 Research