Loading the SOTA2 catalog…
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model · SOTA2 Research