Loading the SOTA2 catalog…
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization · SOTA2 Research