Loading the SOTA2 catalog…
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token · SOTA2 Research