Loading the SOTA2 catalog…
M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning · SOTA2 Research