Loading the SOTA2 catalog…
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers · SOTA2 Research