Loading the SOTA2 catalog…
VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object Detectio · SOTA2 Research