Loading the SOTA2 catalog…
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model · SOTA2 Research