Loading the SOTA2 catalog…
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding · SOTA2 Research