Loading the SOTA2 catalog…
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs · SOTA2 Research