Loading the SOTA2 catalog…
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models · SOTA2 Research