Text-to-image generation on DPG-Bench
88.79Average ScoreNucleus-Image
Evaluation Results
| Method | Links | |
|---|---|---|
| Nucleus-ImageConfiguration=1024 × 10242026.04 | 88.79 | |
| X-OmniModel category=Unified Multimodal Models2026.03 | 87.65 | |
| SD 3.5-ft-10k + MixFlowSampling steps=402025.12 | 86.16 | |
| Show-o2Model category=Unified Multimodal Models2026.03 | 86.14 | |
| BAGEL + UNOModel Type=Unified2026.05 | 86.12 | |
| UniComModel category=Unified Multimodal Models2026.03 | 85.92 | |
| SD 3.5-ft-10k + MixFlowSampling steps=102025.12 | 85.38 | |
| BAGELModel category=Unified Multimodal Models, LLM rewriters=true2026.03 | 85.07 | |
| SD 3.5-ft-20kSampling steps=402025.12 | 84.9 | |
| SD 3.5-ft-10kSampling steps=402025.12 | 84.87 | |
| SD 3.5Sampling steps=402025.12 | 84.8 | |
| LI-DiT-10BModel=LI-DiT-10B, Parameters=10B2024.06 | 84.6 | |
| OneCAT-3BModel Type=Unified2026.05 | 84.53 | |
| MogaoModel category=Unified Multimodal Models2026.03 | 84.33 | |
| Mogao-7BModel Type=Unified2026.05 | 84.33 | |
| SpatialFusionCategory=Unified Generative Model2026.04 | 84.28 | |
| TarModel category=Unified Multimodal Models2026.03 | 84.19 | |
| Janus-ProModel category=Unified Multimodal Models2026.03 | 84.19 | |
| Janus-Pro-7BModel Type=Unified2026.05 | 84.19 | |
| Janus-ProCategory=Unified Generative Model2026.04 | 84.17 | |
| SD3Category=Expertise Generative Model2026.04 | 84.08 | |
| BAGELModel Type=Unified2026.05 | 84.03 | |
| FLUX.1 [Dev]Model category=Generation-only Models2026.03 | 84 | |
| FLUX.1-devDiffusion Params=12B, Resolution=512×5122025.11 | 84 | |
| FLUX.1-devModel Type=Gen. Only2026.05 | 84 | |
| SD-3.5-LargeModel Type=Gen. Only, Evaluation Source=locally evaluated2026.05 | 83.86 | |
| FLUX.1-devCategory=Expertise Generative Model2026.04 | 83.79 | |
| SD-3.5-MediumModel Type=Gen. Only, Evaluation Source=locally evaluated2026.05 | 83.79 | |
| OmniGen2Diffusion Params=4B, Resolution=512×5122025.11 | 83.6 | |
| OmniGen2Model category=Unified Multimodal Models, LLM rewriters=true2026.03 | 83.57 | |
| OmniGen2Category=Unified Generative Model2026.04 | 83.57 | |
| OmniGen2Model Type=Gen. Only2026.05 | 83.57 | |
| DALL-E 3Text Pretrain=Flan-T5-XXL, Res.=1024, rewriting=true2024.12 | 83.5 | |
| DALL-E 3Model=DALL-E 32024.06 | 83.5 | |
| DALL-E 3Resolution=512×5122025.11 | 83.5 | |
| DALL-E 3Category=Expertise Generative Model2026.04 | 83.5 | |
| InfinityModel Type=Gen. Only2026.05 | 83.46 | |
| RPiAEDDT Head usage=with DDT Head2026.03 | 83.34 | |
| RPiAEDDT Head usage=not included2026.03 | 82.59 | |
| Ming-UniVisionModel category=Unified Multimodal Models2026.03 | 82.12 | |
| SD 3.5-ft-20kSampling steps=102025.12 | 81.83 | |
| SD 3.5-ft-10kSampling steps=102025.12 | 81.8 | |
| SD 3.5Sampling steps=102025.12 | 81.78 | |
| LI-DiT-1BModel=LI-DiT-1B, Parameters=1B2024.06 | 81.65 | |
| Emu3Model category=Unified Multimodal Models2026.03 | 81.6 | |
| BLIP3-o-8BModel Type=Unified2026.05 | 81.6 | |
| DeCo-XXL/16Diffusion Params=1.1B, Resolution=512×5122025.11 | 81.4 | |
| UniWorld-V1Category=Unified Generative Model2026.04 | 81.38 | |
| PixNerd-XXL/16Diffusion Params=1.2B, Resolution=512×5122025.11 | 80.9 | |
| RAE-DINOv2-BDDT Head usage=with DDT Head2026.03 | 80.82 | |
| RAE-DINOv2-SDDT Head usage=not included2026.03 | 80.81 | |
| EMU3Res.=512, #Steps=40962024.12 | 80.6 | |
| Emu3-8BModel Type=Unified2026.05 | 80.6 | |
| VAVAEDDT Head usage=with DDT Head2026.03 | 80.32 | |
| JanusModel Type=Unified2026.05 | 79.68 | |
| BLIP3oDiffusion Params=4B, Resolution=512×5122025.11 | 79.4 | |
| SD-XLCategory=Expertise Generative Model2026.04 | 79.26 | |
| VAVAEDDT Head usage=not included2026.03 | 78.69 | |
| RAE-DINOv2-BDDT Head usage=not included2026.03 | 78.58 | |
| Flux-VAEDDT Head usage=not included2026.03 | 77.44 | |
| SDXLText Pretrain=CLIP ViT-bigG, Res.=1024, #Steps=402024.12 | 74.65 | |
| SD XLModel=SD XL2024.06 | 74.65 | |
| SDXLModel Type=Gen. Only2026.05 | 74.65 | |
| TokenFlowRes.=256, #Steps=252024.12 | 73.38 | |
| PixArt-alphaText Pretrain=Flan-T5-XXL, Res.=512, #Steps=202024.12 | 71.11 | |
| PixArt-αModel=PixArt-α2024.06 | 71.11 | |
| PixArt-alphaCategory=Expertise Generative Model2026.04 | 71.11 | |
| PixArt-αDiffusion Params=0.6B, Resolution=512×5122025.11 | 71.1 | |
| VARRes.=256, #Steps=282024.12 | 71.08 | |
| SD v2Model=SD v22024.06 | 68.09 | |
| Show-oText Pretrain=Phi-1.5, Res.=256, #Steps=162024.12 | 67.27 | |
| Show-oCategory=Unified Generative Model2026.04 | 67.27 | |
| LlamaGenText Pretrain=Flan-T5-XL, Res.=512, #Steps=10242024.12 | 64.84 | |
| SD v1.5Text Pretrain=CLIP ViT-L/14, Res.=512, #Steps=502024.12 | 63.18 | |
| SD v1.5Model=SD v1.52024.06 | 63.18 | |
| VFM-VAE + BLIP3-oResolution=256px, Pre-training epoch=12025.10 | 59.1 | |
| VA-VAE + BLIP3-oResolution=256px, Pre-training epoch=12025.10 | 55.4 |