Text-to-image generation on MS-COCO (val)
1.53FIDZEUS-Medium
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ZEUS-MediumModel=FLUX, Scheduler=Euler (Flow)2026.04 | 1.53 | — | — | — | — | — | 30.19 | 0.047 | 2.09 | |
| SADAModel=FLUX, Scheduler=Euler (Flow)2026.04 | 1.95 | — | — | — | — | — | 29.44 | 0.06 | 2.02 | |
| ZEUS-FastModel=FLUX, Scheduler=Euler (Flow)2026.04 | 2.49 | — | — | — | — | — | 26.77 | 0.079 | 2.47 | |
| PartiModel type=Autoregressive, Number of parameters=20B, Finetuned=true2022.11 | 3.22 | — | — | — | — | — | — | — | — | |
| SADAModel=SDXL, Scheduler=DPM++2026.04 | 3.51 | — | — | — | — | — | 29.36 | 0.084 | 1.86 | |
| ZEUS-MediumModel=SDXL, Scheduler=DPM++2026.04 | 3.59 | — | — | — | — | — | 29.17 | 0.084 | 1.87 | |
| SADAModel=SDXL, Scheduler=Euler2026.04 | 3.76 | — | — | — | — | — | 28.97 | 0.093 | 1.85 | |
| ZEUS-MediumModel=SDXL, Scheduler=Euler2026.04 | 3.87 | — | — | — | — | — | 28.66 | 0.095 | 1.85 | |
| SADAModel=SD-2, Scheduler=DPM++2026.04 | 4.02 | — | — | — | — | — | 26.34 | 0.094 | 1.8 | |
| SADAModel=SD-2, Scheduler=Euler2026.04 | 4.26 | — | — | — | — | — | 26.25 | 0.1 | 1.81 | |
| AdaptiveDiffusionModel=SD-2, Scheduler=DPM++2026.04 | 4.35 | — | — | — | — | — | 24.3 | 0.1 | 1.45 | |
| ZEUS-MediumModel=SD-2, Scheduler=DPM++2026.04 | 4.46 | — | — | — | — | — | 25.72 | 0.104 | 1.85 | |
| ZEUS-TurboModel=FLUX, Scheduler=Euler (Flow)2026.04 | 4.52 | — | — | — | — | — | 21.8 | 0.171 | 3.22 | |
| VeCoRSolver=SDE (E-M), CFG Scale (ω)=2.0, Steps=50, Augmentation Strategy=Random Crop and Resize (RCR)2025.11 | 4.55 | — | — | — | — | — | — | — | — | |
| AdaptiveDiffusionModel=SDXL, Scheduler=DPM++2026.04 | 4.59 | — | — | — | — | — | 26.16 | 0.125 | 1.65 | |
| MMDiT + REPABackbone=MMDiT, Representation Learner=REPA2026.01 | 4.6 | — | — | 20.9 | — | — | — | — | — | |
| MMDiT + VAE-REPABackbone=MMDiT, Representation Learner=VAE-REPA2026.01 | 4.67 | — | — | 20.92 | — | — | — | — | — | |
| MMDiT + SRABackbone=MMDiT, Representation Learner=SRA2026.01 | 4.74 | — | — | 21.12 | — | — | — | — | — | |
| ΔFMSolver=SDE (E-M), CFG Scale (ω)=2.0, Steps=502025.11 | 4.78 | — | — | — | — | — | — | — | — | |
| VeCoRSolver=ODE (Heun), CFG Scale (ω)=2.0, Steps=50, Augmentation Strategy=Random Crop and Resize (RCR)2025.11 | 4.82 | — | — | — | — | — | — | — | — | |
| TeaCacheModel=FLUX, Scheduler=Euler (Flow)2026.04 | 4.89 | — | — | — | — | — | 19.14 | 0.216 | 2 | |
| MMDiT + REPASolver=ODE (Heun), CFG Scale (ω)=2.0, Steps=502025.11 | 5.03 | — | — | — | — | — | — | — | — | |
| VeCoRSolver=SDE (E-M), CFG Scale (ω)=2.0, Steps=50, Augmentation Strategy=Random Channel Shuffle (RCS)2025.11 | 5.03 | — | — | — | — | — | — | — | — | |
| ZEUS-MediumModel=SD-2, Scheduler=Euler2026.04 | 5.06 | — | — | — | — | — | 25.37 | 0.118 | 1.86 | |
| MMDiT BaselineBackbone=MMDiT2026.01 | 5.08 | — | — | 20.54 | — | — | — | — | — | |
| ΔFMSolver=ODE (Heun), CFG Scale (ω)=2.0, Steps=502025.11 | 5.16 | — | — | — | — | — | — | — | — | |
| Re-ImagenModel type=Diffusion, Finetuned=true2022.11 | 5.25 | — | — | — | — | — | — | — | — | |
| MMDiTType=Diffusion2025.12 | 5.3 | — | — | — | — | — | — | — | — | |
| VeCoRSolver=ODE (Heun), CFG Scale (ω)=2.0, Steps=50, Augmentation Strategy=Random Channel Shuffle (RCS)2025.11 | 5.3 | — | — | — | — | — | — | — | — | |
| ZEUS-FastModel=SDXL, Scheduler=DPM++2026.04 | 5.39 | — | — | — | — | — | 26.38 | 0.129 | 1.93 | |
| U-ViT/S/2 (Deep)Type=Diffusion2025.12 | 5.45 | — | — | — | — | — | — | — | — | |
| U-ViT-S/2Type=Diffusion2025.12 | 5.95 | — | — | — | — | — | — | — | — | |
| MMDiT + REPASolver=SDE (E-M), CFG Scale (ω)=2.0, Steps=502025.11 | 6.03 | — | — | — | — | — | — | — | — | |
| AdaptiveDiffusionModel=SDXL, Scheduler=Euler2026.04 | 6.11 | — | — | — | — | — | 24.33 | 0.168 | 2.01 | |
| CONPREDIFF_conModel Type=Continuous Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 6.21 | — | — | — | — | — | — | — | — | |
| APEXNFEs=2, Throughput (samples/s)=5.72, Latency (s)=0.23, Params (B)=1.6, Training Protocol=Full tuning2026.04 | 6.42 | 28.24 | — | — | — | — | — | — | — | |
| APEXNFEs=2, Throughput (samples/s)=3.30, Latency (s)=0.45, Params (B)=20, Training Protocol=Full tuning2026.04 | 6.44 | 28.51 | — | — | — | — | — | — | — | |
| ZEUS-FastModel=SDXL, Scheduler=Euler2026.04 | 6.47 | — | — | — | — | — | 25.15 | 0.153 | 1.93 | |
| Sana-SprintNFEs=2, Throughput (samples/s)=5.68, Latency (s)=0.24, Params (B)=1.6, Training Protocol=Full tuning2026.04 | 6.5 | 28.45 | — | — | — | — | — | — | — | |
| APEXNFEs=2, Throughput (samples/s)=3.17, Latency (s)=0.47, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 6.51 | 28.42 | — | — | — | — | — | — | — | |
| Sana-SprintNFEs=2, Throughput (samples/s)=6.46, Latency (s)=0.25, Params (B)=0.6, Training Protocol=Full tuning2026.04 | 6.54 | 28.4 | — | — | — | — | — | — | — | |
| ΔFMSolver=SDE (E-M), CFG Scale (ω)=1.0, Steps=502025.11 | 6.64 | — | — | — | — | — | — | — | — | |
| GroupDiff-4Type=Diffusion, Base Model=DiT-XL/2 w/ Cross-Attention2025.12 | 6.65 | — | — | — | — | — | — | — | — | |
| VeCoRSolver=SDE (E-M), CFG Scale (ω)=1.0, Steps=50, Augmentation Strategy=Random Channel Shuffle (RCS)2025.11 | 6.65 | — | — | — | — | — | — | — | — | |
| CONPREDIFF_disModel Type=Discrete Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 6.67 | — | — | — | — | — | — | — | — | |
| APEXNFEs=2, Throughput (samples/s)=3.21, Latency (s)=0.49, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=32, Trainable Params (B)=0.22026.04 | 6.72 | 28.71 | — | — | — | — | — | — | — | |
| TwinFlowNFEs=2, Throughput (samples/s)=3.15, Latency (s)=0.48, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 6.73 | 28.57 | — | — | — | — | — | — | — | |
| Ernie-ViLG 2.0num of model=22025.12 | 6.75 | — | — | — | — | — | — | — | — | |
| APEXNFEs=2, Throughput (samples/s)=6.50, Latency (s)=0.25, Params (B)=0.6, Training Protocol=Full tuning2026.04 | 6.75 | 28.33 | — | — | — | — | — | — | — | |
| Qwen-Image-LightningNFEs=2, Throughput (samples/s)=3.15, Latency (s)=0.48, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 6.76 | 28.37 | — | — | — | — | — | — | — | |
| APEXNFEs=1, Throughput (samples/s)=6.84, Latency (s)=0.20, Params (B)=1.6, Training Protocol=Full tuning2026.04 | 6.78 | 28.12 | — | — | — | — | — | — | — | |
| RCGMNFEs=2, Throughput (samples/s)=3.15, Latency (s)=0.48, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 6.8 | 28.63 | — | — | — | — | — | — | — | |
| APEXNFEs=1, Throughput (samples/s)=3.50, Latency (s)=0.34, Params (B)=20, Training Protocol=Full tuning2026.04 | 6.87 | 28.66 | — | — | — | — | — | — | — | |
| Re-ImagenZero-shot=true2022.11 | 6.88 | — | — | — | — | — | — | — | — | |
| Re-ImagenModel type=Diffusion, Finetuned=false2022.11 | 6.88 | — | — | — | — | — | — | — | — | |
| eDiff-IModel Type=Continuous Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 6.95 | — | — | — | — | — | — | — | — | |
| e-diffnum of model=32025.12 | 6.95 | — | — | — | — | 5.5 | — | — | — | |
| DiT-XL/2 w/ Cross-AttentionType=Diffusion2025.12 | 6.95 | — | — | — | — | — | — | — | — | |
| APEXNFEs=1, Throughput (samples/s)=7.30, Latency (s)=0.20, Params (B)=0.6, Training Protocol=Full tuning2026.04 | 6.99 | 28.36 | — | — | — | — | — | — | — | |
| Sana-SprintNFEs=1, Throughput (samples/s)=7.22, Latency (s)=0.21, Params (B)=0.6, Training Protocol=Full tuning2026.04 | 7.04 | 28.04 | — | — | — | — | — | — | — | |
| Qwen-Image-LightningNFEs=1, Throughput (samples/s)=3.29, Latency (s)=0.40, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 7.06 | 28.35 | — | — | — | — | — | — | — | |
| SDXL-DMD2NFEs=1, Throughput (samples/s)=3.36, Latency (s)=0.32, Params (B)=0.9, Training Protocol=Full tuning, Model Strategy=Distinct models per NFE2026.04 | 7.1 | 28.93 | — | — | — | — | — | — | — | |
| APEXNFEs=1, Throughput (samples/s)=3.27, Latency (s)=0.39, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 7.14 | 28.45 | — | — | — | — | — | — | — | |
| APEXNFEs=1, Throughput (samples/s)=3.29, Latency (s)=0.39, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=32, Trainable Params (B)=0.22026.04 | 7.22 | 28.62 | — | — | — | — | — | — | — | |
| PartiZero-shot=true2022.11 | 7.23 | — | — | — | — | — | — | — | — | |
| PartiModel type=Autoregressive, Number of parameters=20B, Finetuned=false2022.11 | 7.23 | — | — | — | — | — | — | — | — | |
| Parti-20BLM=true, MG=false, FIG=false, Zero-shot=true2023.09 | 7.23 | — | — | — | — | — | — | — | — | |
| Partimode=zero-shot, samples=30k, model_category=unimodal generation models2023.07 | 7.23 | — | — | — | — | — | — | — | — | |
| PartiModel Type=Autoregressive, Zero-shot=true, Resolution=256x2562024.01 | 7.23 | — | — | — | — | — | — | — | — | |
| FLUX-schnellNFEs=1, Throughput (samples/s)=1.58, Latency (s)=0.68, Params (B)=12.0, Training Protocol=Full tuning2026.04 | 7.26 | 28.49 | — | — | — | — | — | — | — | |
| ImagenZero-shot=true2022.11 | 7.27 | — | — | — | — | — | — | — | — | |
| ImagenModel type=Diffusion, Finetuned=false2022.11 | 7.27 | — | — | — | — | — | — | — | — | |
| Imagen-3.4BLM=true, MG=false, FIG=false, Zero-shot=true2023.09 | 7.27 | — | — | — | — | — | — | — | — | |
| Imagenmode=zero-shot, samples=30k, model_category=unimodal generation models2023.07 | 7.27 | — | — | — | — | — | — | — | — | |
| ImagenModel Type=Continuous Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 7.27 | — | — | — | — | — | — | — | — | |
| U-NetType=Diffusion2025.12 | 7.32 | — | — | — | — | — | — | — | — | |
| TwinFlowNFEs=1, Throughput (samples/s)=3.29, Latency (s)=0.40, Backbone Params (B)=20, Training Protocol=LoRA, LoRA rank=64, Trainable Params (B)=0.42026.04 | 7.32 | 28.29 | — | — | — | — | — | — | — | |
| DeepCacheModel=SDXL, Scheduler=Euler2026.04 | 7.36 | — | — | — | — | — | 22 | 0.223 | 2.16 | |
| DeepCacheModel=SD-2, Scheduler=Euler2026.04 | 7.4 | — | — | — | — | — | 18.91 | 0.239 | 1.45 | |
| Make-A-SceneModel Type=Autoregressive, Zero-shot=false, Resolution=256x2562024.01 | 7.55 | — | — | — | — | — | — | — | — | |
| AdaptiveDiffusionModel=SD-2, Scheduler=Euler2026.04 | 7.58 | — | — | — | — | — | 21.94 | 0.173 | 1.89 | |
| SDXL-DMD2NFEs=2, Throughput (samples/s)=2.89, Latency (s)=0.40, Params (B)=0.9, Training Protocol=Full tuning, Model Strategy=Distinct models per NFE2026.04 | 7.61 | 28.87 | — | — | — | — | — | — | — | |
| Sana-SprintNFEs=1, Throughput (samples/s)=6.71, Latency (s)=0.21, Params (B)=1.6, Training Protocol=Full tuning2026.04 | 7.69 | 28.27 | — | — | — | — | — | — | — | |
| FLUX-schnellNFEs=2, Throughput (samples/s)=0.92, Latency (s)=1.15, Params (B)=12.0, Training Protocol=Full tuning2026.04 | 7.75 | 28.25 | — | — | — | — | — | — | — | |
| DeepCacheModel=SD-2, Scheduler=DPM++2026.04 | 7.83 | — | — | — | — | — | 17.7 | 0.271 | 1.43 | |
| Muse-3BLM=true, MG=false, FIG=false, Zero-shot=true2023.09 | 7.88 | — | — | — | — | — | — | — | — | |
| MuseModel Type=Non-Autoregressive, Zero-shot=true, Resolution=256x2562024.01 | 7.88 | — | — | — | — | — | — | — | — | |
| VeCoRSolver=SDE (E-M), CFG Scale (ω)=1.0, Steps=50, Augmentation Strategy=Random Crop and Resize (RCR)2025.11 | 7.95 | — | — | — | — | — | — | — | — | |
| LafiteSupervised=true2022.11 | 8.12 | — | — | — | — | — | — | — | — | |
| LAFITEModel Type=GAN, Zero-shot=false, Resolution=256x2562024.01 | 8.12 | — | — | — | — | — | — | — | — | |
| LAFITEType=GAN2025.12 | 8.12 | — | — | — | — | — | — | — | — | |
| Simple DiffusionModel Type=Continuous Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 8.32 | — | — | — | — | — | — | — | — | |
| Improved VQ-DiffusionModel Type=Discrete Diffusion, Zero-shot=true, Resolution=256x2562024.01 | 8.44 | — | — | — | — | — | — | — | — | |
| DREAMLLM-7BLM=true, MG=true, FIG=true, Zero-shot=true2023.09 | 8.46 | — | — | — | — | — | — | — | — | |
| DeepCacheModel=SDXL, Scheduler=DPM++2026.04 | 8.48 | — | — | — | — | — | 21.35 | 0.255 | 1.74 | |
| DREAMLLM-7BLM=true, MG=true, FIG=true, Zero-shot=true, stage=stage I alignment training2023.09 | 8.76 | — | — | — | — | — | — | — | — | |
| ToCaModel=FLUX, Scheduler=Euler (Flow)2026.04 | 8.84 | — | — | — | — | — | 17.7 | 0.352 | 1.52 | |
| FridoType=Diffusion2025.12 | 8.97 | — | — | — | — | — | — | — | — | |
| Teacher modelInference Steps=20 steps, Resolution=512x512, Batch size=12024.03 | 9.273 | 0.2863 | 1.44 | — | — | — | — | — | — | |
| XMC-GANSupervised=true2022.11 | 9.33 | — | — | — | — | — | — | — | — |