Loading the SOTA2 catalog…
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability · SOTA2 Research