Loading the SOTA2 catalog…
ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World Data · SOTA2 Research