Loading the SOTA2 catalog…
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model · SOTA2 Research