Loading the SOTA2 catalog…
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond · SOTA2 Research