Loading the SOTA2 catalog…
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models · SOTA2 Research