Loading the SOTA2 catalog…
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models · SOTA2 Research