Loading the SOTA2 catalog…
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation · SOTA2 Research