Text-to-image synthesis on CUB-200-2011 (test)
10.32FIDVQ-Diffusion
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| VQ-Diffusionvariant=Full2021.11 | 10.32 | — | — | — | — | — | — | |
| EFF-T2I2021.11 | 11.17 | — | — | — | — | — | — | |
| VQ-Diffusionvariant=Base2021.11 | 11.94 | — | — | — | — | — | — | |
| VQ-Diffusionvariant=Small2021.11 | 12.97 | — | — | — | — | — | — | |
| DM-GAN + CLcontrastive_learning=true2021.07 | 14.38 | — | — | — | 4.77 | 78.99 | — | |
| DF-GAN2021.11 | 14.81 | — | — | — | — | — | — | |
| DM-GANpre-trained baseline=true2021.07 | 15.1 | — | — | — | 4.66 | 75.86 | — | |
| DAE-GAN2021.11 | 15.19 | — | — | — | — | — | — | |
| StackGAN++2021.11 | 15.3 | — | — | — | — | — | — | |
| DM-GAN2021.11 | 16.09 | — | — | — | — | — | — | |
| AttnGAN + CLapproach=contrastive learning2021.07 | 16.34 | — | — | — | 4.42 | 69.64 | — | |
| SEGAN2021.11 | 18.17 | — | — | — | — | — | — | |
| AttnGANpublic_pre-trained_model=true2021.07 | 20.85 | — | — | — | 4.33 | 67.09 | — | |
| AttnGAN2021.11 | 23.98 | — | — | — | — | — | — | |
| MSGANConditioning setting=Conditioned on text descriptions2019.03 | 25.53 | 30.6 | 0.073 | 0.373 | — | — | — | |
| StackGAN++Conditioning setting=Conditioned on text descriptions2019.03 | 25.99 | 38.2 | 0.092 | 0.362 | — | — | — | |
| StackGAN++Conditioning setting=Conditioned on text codes2019.03 | 27.12 | 39 | 0.102 | 0.156 | — | — | — | |
| MSGANConditioning setting=Conditioned on text codes2019.03 | 27.94 | 30.6 | 0.095 | 0.207 | — | — | — | |
| StackGAN2021.11 | 51.89 | — | — | — | — | — | — | |
| DALL-E2021.11 | 56.1 | — | — | — | — | — | — | |
| Gumbel-VQgenerative model=diffusion, resolution=256x2562023.03 | — | — | — | — | — | — | 16.93 | |
| Reg-VQgenerative model=diffusion, resolution=256x2562023.03 | — | — | — | — | — | — | 14.14 | |
| VQ-GANgenerative model=diffusion, resolution=256x2562023.03 | — | — | — | — | — | — | 17.43 | |
| VQ-VAEgenerative model=diffusion, resolution=256x2562023.03 | — | — | — | — | — | — | 26.32 |