Loading the SOTA2 catalog…
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models · SOTA2 Research