Loading the SOTA2 catalog…
GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection · SOTA2 Research