Loading the SOTA2 catalog…
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models · SOTA2 Research