Loading the SOTA2 catalog…
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation · SOTA2 Research