Loading the SOTA2 catalog…
EVLM: An Efficient Vision-Language Model for Visual Understanding · SOTA2 Research