Loading the SOTA2 catalog…
F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models · SOTA2 Research