Loading the SOTA2 catalog…
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models · SOTA2 Research