Loading the SOTA2 catalog…
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion · SOTA2 Research