Loading the SOTA2 catalog…
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks · SOTA2 Research