Loading the SOTA2 catalog…
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios · SOTA2 Research