Loading the SOTA2 catalog…
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models · SOTA2 Research