Loading the SOTA2 catalog…
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models · SOTA2 Research