Loading the SOTA2 catalog…
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection · SOTA2 Research