Loading the SOTA2 catalog…
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention · SOTA2 Research