Loading the SOTA2 catalog…
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models · SOTA2 Research