Loading the SOTA2 catalog…
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity · SOTA2 Research