Loading the SOTA2 catalog…
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? · SOTA2 Research