Conditional Image Generation on COCO (val)
0.608CLIP-TOpen unCLIP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Open unCLIPTraining Strategy=Training from scratch, Reusable to custom models=false, Compatible with controllable tools=false, Multimodal prompts=false, Trainable parameters=893M2023.08 | 0.608 | 0.858 | |
| Kandinsky 2.1Training Strategy=Training from scratch, Reusable to custom models=false, Compatible with controllable tools=false, Multimodal prompts=false, Trainable parameters=1229M2023.08 | 0.599 | 0.855 | |
| IP-AdapterTraining Strategy=Adapters, Reusable to custom models=true, Compatible with controllable tools=true, Multimodal prompts=true, Trainable parameters=22M2023.08 | 0.588 | 0.828 | |
| Versatile DiffusionTraining Strategy=Training from scratch, Reusable to custom models=false, Compatible with controllable tools=false, Multimodal prompts=true, Trainable parameters=860M2023.08 | 0.587 | 0.83 | |
| SD unCLIPTraining Strategy=Fine-tuning from text-to-image model, Reusable to custom models=false, Compatible with controllable tools=false, Multimodal prompts=false, Trainable parameters=870M2023.08 | 0.584 | 0.81 | |
| SD Image VariationsTraining Strategy=Fine-tuning from text-to-image model, Reusable to custom models=false, Compatible with controllable tools=false, Multimodal prompts=false, Trainable parameters=860M2023.08 | 0.548 | 0.76 | |
| Uni-ControlNet (Global Control)Training Strategy=Adapters, Reusable to custom models=true, Compatible with controllable tools=true, Multimodal prompts=true, Trainable parameters=47M2023.08 | 0.506 | 0.736 | |
| T2I-Adapter (Style)Training Strategy=Adapters, Reusable to custom models=true, Compatible with controllable tools=true, Multimodal prompts=true, Trainable parameters=39M2023.08 | 0.485 | 0.648 | |
| ControlNet ShuffleTraining Strategy=Adapters, Reusable to custom models=true, Compatible with controllable tools=true, Multimodal prompts=true, Trainable parameters=361M2023.08 | 0.421 | 0.616 |