Loading the SOTA2 catalog…
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation · SOTA2 Research