Loading the SOTA2 catalog…
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP · SOTA2 Research