Loading the SOTA2 catalog…
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A · SOTA2 Research