Loading the SOTA2 catalog…
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification · SOTA2 Research