Loading the SOTA2 catalog…
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation · SOTA2 Research