Loading the SOTA2 catalog…
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization · SOTA2 Research