Text-to-Image Generation on PartiPrompts (HPS/ImageReward Suite)
56CLIPScoreTuned-CFG+PG-MAP†
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Tuned-CFG+PG-MAP†Source=Ours, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 56 | 66 | 60.2 | — | 53.6 | — | — | |
| Tuned-CFG+PG-MAP†Source=Ours, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 52.8 | 64.6 | 56.5 | — | 51.3 | — | — | |
| Tuned-CFG*Source=Compare, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 52.1 | 52.7 | 56.4 | — | 47.2 | — | — | |
| Reward-zSource=Ours, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 51.3 | 54.2 | 54.9 | — | 57.4 | — | — | |
| MAP-cSource=Ours, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 51 | 51 | 44.9 | — | 51.6 | — | — | |
| UG*Source=Compare, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 50.7 | 46.9 | 51.4 | — | 46.3 | — | — | |
| PG-MAP† (default)Source=Ours, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 50.6 | 52.8 | 54 | — | 56.8 | — | — | |
| Baseline (reference)Source=-, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 50 | 50 | 50 | — | 50 | — | — | |
| Baseline (reference)Source=-, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 50 | 50 | 50 | — | 50 | — | — | |
| Tuned-CFG*Source=Compare, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 50 | 58.5 | 52.4 | — | 48.2 | — | — | |
| Reward-zSource=Ours, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 49.7 | 47.9 | 56.7 | — | 55.4 | — | — | |
| MAP-cz (λ=0, reward-free)Source=Ours, Backbone=Stable Diffusion 1.5, Sampling Steps=30, Sampling Algorithm=DDIM, CFG Scale=7.52026.06 | 49.5 | 52.6 | 54.9 | — | 56.5 | — | — | |
| MAP-cz (λ=0, reward-free)Source=Ours, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 48.8 | 47.5 | 55.6 | — | 56.7 | — | — | |
| MAP-cSource=Ours, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 48.5 | 50.3 | 49.8 | — | 51.4 | — | — | |
| PG-MAP† (default)Source=Ours, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 48.1 | 47.1 | 56.2 | — | 56.4 | — | — | |
| UG*Source=Compare, Backbone=SDXL, Sampling Steps=50, Sampling Algorithm=DDIM, CFG Scale=5.02026.06 | 47.9 | 50.5 | 51.1 | — | 48.6 | — | — | |
| TaylorSeerCaching=TaylorSeer, Setting=N = 5, O = 2, TMACs=357.39*, s/img (speedup)=7.20 (2.54x)*2025.06 | 32.28 | — | — | 0.94 | — | — | — | |
| ECADModel=FLUX.1-dev, Caching=Ours, Setting=fastest, TMACs=43.60, Latency (ms/img)=778.17 (3.37x)2025.06 | 32.27 | — | — | 0.89 | — | — | — | |
| ECADModel=FLUX.1-dev, Caching=Ours, Setting=fast, TMACs=63.02, Latency (ms/img)=1016.59 (2.58x)2025.06 | 32.24 | — | — | 1.04 | — | — | — | |
| ToCaCaching=ToCa, Setting=N = 4, R = 90%, TMACs=300.41*, s/img (speedup)=7.42 (2.47x)*2025.06 | 32.05 | — | — | 1.09 | — | — | — | |
| FORAModel=PixArt-α, Caching=FORA, Setting=N = 2, TMACs=2.87, Latency (ms/img)=100.57 (1.65x)2025.06 | 32.03 | — | — | 0.91 | — | — | — | |
| NoneModel=PixArt-α, Caching=None, TMACs=5.71, Latency (ms/img)=165.74 (1.00x)2025.06 | 32.01 | — | — | 0.97 | — | — | — | |
| NoneCaching=None, Setting=None, TMACs=1190.25, s/img (speedup)=18.30 (1.00x)2025.06 | 31.98 | — | — | 1.14 | — | — | — | |
| DiCacheModel=FLUX.1-dev, Caching=DiCache, TMACs=62.23, Latency (ms/img)=1161.86 (2.26x)2025.06 | 31.97 | — | — | 0.97 | — | — | — | |
| FORAModel=PixArt-α, Caching=FORA, Setting=N = 3, TMACs=2.02, Latency (ms/img)=82.55 (2.01x)2025.06 | 31.94 | — | — | 0.83 | — | — | — | |
| ECADModel=PixArt-α, Caching=Ours, Setting=fast, TMACs=2.13, Latency (ms/img)=84.09 (1.97x)2025.06 | 31.94 | — | — | 0.99 | — | — | — | |
| FORAModel=PixArt-Σ, Caching=FORA, Setting=N = 3, TMACs=2.02, Latency (ms/img)=82.12 (2.04x)2025.06 | 31.91 | — | — | 0.81 | — | — | — | |
| NoneModel=PixArt-Σ, Caching=None, TMACs=5.71, Latency (ms/img)=167.62 (1.00x)2025.06 | 31.9 | — | — | 1.08 | — | — | — | |
| NoneModel=FLUX.1-dev, Caching=None, TMACs=198.69, Latency (ms/img)=2620.09 (1.00x)2025.06 | 31.88 | — | — | 1.04 | — | — | — | |
| FORAModel=FLUX.1-dev, Caching=FORA, Setting=N = 3, TMACs=69.80, Latency (ms/img)=1073.70 (2.44x)2025.06 | 31.88 | — | — | 0.93 | — | — | — | |
| ECADCaching=Ours, Setting=fast 256→1024, TMACs=376.62, s/img (speedup)=6.96 (2.63x)2025.06 | 31.88 | — | — | 1.05 | — | — | — | |
| ECADModel=PixArt-Σ, Caching=Ours, Setting=fast, TMACs=1.91, Latency (ms/img)=84.84 (1.98x)2025.06 | 31.86 | — | — | 1.02 | — | — | — | |
| ECADCaching=Ours, Setting=slow 256→1024, TMACs=644.05, s/img (speedup)=10.59 (1.73x)2025.06 | 31.82 | — | — | 1.05 | — | — | — | |
| ToCaModel=FLUX.1-dev, Caching=ToCa, Setting=N = 4, R = 90%, TMACs=42.96*, Latency (ms/img)=1576.97 (1.66x)*2025.06 | 31.81 | — | — | 0.93 | — | — | — | |
| TGATEModel=PixArt-α, Caching=TGATE, Setting=m = 15, k = 1, TMACs=4.86, Latency (ms/img)=144.77 (1.14x)2025.06 | 31.7 | — | — | 0.87 | — | — | — | |
| DuCaModel=PixArt-α, Caching=DuCa, Setting=N = 3, R = 60%, TMACs=3.20, Latency (ms/img)=72.53 (2.29x)*2025.06 | 31.53 | — | — | 0.79 | — | — | — | |
| ECADModel=PixArt-α, Caching=Ours, Setting=fastest, TMACs=1.18, Latency (ms/img)=64.24 (2.58x)2025.06 | 31.53 | — | — | 0.77 | — | — | — | |
| ToCaModel=PixArt-α, Caching=ToCa, Setting=N = 3, R = 60%, TMACs=3.17*, Latency (ms/img)=90.71 (1.83x)*2025.06 | 31.46 | — | — | 0.76 | — | — | — | |
| ECADModel=PixArt-α, Caching=Ours, Setting=faster, TMACs=1.46, Latency (ms/img)=69.17 (2.40x)2025.06 | 31.44 | — | — | 0.88 | — | — | — | |
| DuCaModel=PixArt-α, Caching=DuCa, Setting=N = 3, R = 90%, TMACs=2.30, Latency (ms/img)=64.08 (2.59x)*2025.06 | 31.42 | — | — | 0.74 | — | — | — | |
| NoneCaching=None, Setting=40% steps, TMACs=476.10, s/img (speedup)=7.61 (2.41x)2025.06 | 31.38 | — | — | 0.83 | — | — | — | |
| ToCaModel=PixArt-α, Caching=ToCa, Setting=N = 3, R = 90%, TMACs=2.13*, Latency (ms/img)=70.58 (2.35x)*2025.06 | 31.35 | — | — | 0.68 | — | — | — | |
| FORACaching=FORA, Setting=N = 3, TMACs=416.88, s/img (speedup)=7.62 (2.40x)2025.06 | 31.2 | — | — | 0.69 | — | — | — | |
| TaylorSeerModel=FLUX.1-dev, Caching=TaylorSeer, Setting=N = 5, O = 2, TMACs=59.88*, Latency (ms/img)=1028.66 (2.55x)*2025.06 | 31.16 | — | — | 0.54 | — | — | — | |
| ToCaModel=PixArt-Σ, Caching=ToCa†, Setting=N = 3, R = 60%, TMACs=3.17*, Latency (ms/img)=94.28 (1.78x)*2025.06 | 31.03 | — | — | 0.19 | — | — | — | |
| ToCaModel=PixArt-Σ, Caching=ToCa†, Setting=N = 3, R = 90%, TMACs=2.13*, Latency (ms/img)=73.03 (2.30x)*2025.06 | 30.89 | — | — | 0.14 | — | — | — | |
| TaylorSeerModel=FLUX.1-dev, Caching=TaylorSeer, Setting=N = 6, O = 1, TMACs=49.97*, Latency (ms/img)=865.97 (3.03x)*2025.06 | 29.88 | — | — | 0.02 | — | — | — | |
| TGATEModel=PixArt-α, Caching=TGATE, Setting=m = 10, k = 5, TMACs=3.47, Latency (ms/img)=108.52 (1.53x)2025.06 | 28.9 | — | — | -0.27 | — | — | — | |
| DTMHead Type=MLP, Head Size=mid, Head Seq scale=1, Head batch size=16, Weighting=πln(0, 1) × πln(0, 1), Sampling=linear2025.12 | 27 | — | 5.47 | 0.63 | 21.3 | — | — | |
| DTM++Head Type=MLP, Head Size=mid, Head Seq scale=1, Head batch size=16, Weighting=πln(0, 1) × πln(0, 1), Sampling=c = 0.2, τ = 12025.12 | 27 | — | 5.57 | 0.7 | 21.4 | — | — | |
| MARHead Type=MLP, Head Size=mid, Head Seq scale=1, Head batch size=4, Weighting=U(0, 1), Sampling=linear2025.12 | 27 | — | 4.95 | 0.33 | 20.7 | — | — | |
| MARSampling=argmax2025.12 | 26.8 | — | 5.15 | 0.14 | 20.7 | — | — | |
| ARSampling=argmax2025.12 | 26.7 | — | 4.81 | -0.01 | 20.4 | — | — | |
| DTMHead Type=MLP, Head Size=mid, Head Seq scale=1, Head batch size=4, Weighting=U(0, 1) × U(0, 1), Sampling=linear2025.12 | 26.6 | — | 5.46 | 0.51 | 21.2 | — | — | |
| DTMHead Type=Transformer, Head Size=mid, Head Seq scale=4, Head batch size=16, Weighting=πln(0, 1) × πln(0, 1), Sampling=linear2025.12 | 26.6 | — | 5.59 | 0.62 | 21.4 | — | — | |
| DTM+Head Type=Transformer, Head Size=mid, Head Seq scale=4, Head batch size=16, Weighting=πln(0, 1) × πln(0, 1), Sampling=c = 0.8, τ = 322025.12 | 26.6 | — | 5.58 | 0.63 | 21.4 | — | — | |
| FMWeighting=πln(0, 1), Sampling=linear2025.12 | 26.6 | — | 5.44 | 0.48 | 21.2 | — | — | |
| DTMHead Type=Transformer, Head Size=mid, Head Seq scale=1, Head batch size=4, Weighting=U(0, 1) × U(0, 1), Sampling=linear2025.12 | 26.5 | — | 5.54 | 0.51 | 21.3 | — | — | |
| DTMHead Type=Convolution, Head Size=mid, Head Seq scale=1, Head batch size=4, Weighting=U(0, 1) × U(0, 1), Sampling=linear2025.12 | 26.4 | — | 5.52 | 0.51 | 21.3 | — | — | |
| DTM (Dense)Head Type=Dense, Weighting=U(0, 1) × U(0, 1), Sampling=linear2025.12 | 26.1 | — | 5.44 | 0.36 | 21.2 | — | — | |
| FMWeighting=U(0, 1), Sampling=linear2025.12 | 26.1 | — | 5.33 | 0.34 | 21.1 | — | — | |
| ARHead Type=MLP, Head Size=mid, Head Seq scale=1, Head batch size=4, Weighting=U(0, 1), Sampling=linear2025.12 | 24.9 | — | 4.5 | -0.43 | 20.1 | — | — | |
| ArenaPOBackbone=Stable Diffusion XL2026.05 | 0.3671 | 0.2906 | 5.9003 | 1.0672 | 0.2299 | — | — | |
| DSPOBackbone=Stable Diffusion XL2026.05 | 0.3664 | 0.2871 | 5.6947 | 1.0514 | 0.2261 | — | — | |
| Diffusion LAIRBase Model=SDXL, Samples per prompt=52026.05 | 0.366 | 0.292 | 5.845 | 1.104 | 22.765 | — | — | |
| Diff.-DPOBase Model=SDXL, Samples per prompt=52026.05 | 0.365 | 0.2894 | 5.824 | 1.066 | 22.693 | — | — | |
| SDPOBackbone=Stable Diffusion XL2026.05 | 0.3645 | 0.2907 | 5.7882 | 1.0654 | 0.229 | — | — | |
| Diff.-DPOBackbone=Stable Diffusion XL2026.05 | 0.3629 | 0.29 | 5.8294 | 1.0638 | 0.2279 | — | — | |
| Pre-TrainedBackbone=Stable Diffusion XL2026.05 | 0.3591 | 0.288 | 5.7901 | 0.8573 | 0.2277 | — | — | |
| InPOBase Model=SDXL, Samples per prompt=52026.05 | 0.359 | 0.2908 | 5.872 | 1.045 | 22.723 | — | — | |
| MaPOBackbone=Stable Diffusion XL2026.05 | 0.358 | 0.2902 | 5.8921 | 0.9324 | 0.2278 | — | — | |
| SDXLBase Model=SDXL, Samples per prompt=52026.05 | 0.356 | 0.2843 | 5.809 | 0.776 | 22.425 | — | — | |
| SFTBackbone=Stable Diffusion XL2026.05 | 0.3559 | 0.2834 | 5.6496 | 0.7515 | 0.2221 | — | — | |
| MaPOBase Model=SDXL, Samples per prompt=52026.05 | 0.354 | 0.2861 | 5.957 | 0.873 | 22.399 | — | — | |
| Diffusion LAIRBase Model=SD1.5, Samples per prompt=52026.05 | 0.3485 | 0.286 | 5.671 | 0.8107 | 21.992 | — | — | |
| InPOBase Model=SD1.5, Samples per prompt=52026.05 | 0.3457 | 0.2842 | 5.613 | 0.7203 | 21.735 | — | — | |
| ArenaPOBackbone=Stable Diffusion 1.52026.05 | 0.3445 | 0.2833 | 5.6164 | 0.6291 | 0.2198 | — | — | |
| SDPOBackbone=Stable Diffusion 1.52026.05 | 0.3423 | 0.2815 | 5.588 | 0.5425 | 0.2187 | — | — | |
| Diff.-DPOBackbone=Stable Diffusion 1.52026.05 | 0.3391 | 0.2755 | 5.4045 | 0.256 | 0.2167 | — | — | |
| Diff.-KTOBase Model=SD1.5, Samples per prompt=52026.05 | 0.339 | 0.2825 | 5.568 | 0.5941 | 21.55 | — | — | |
| SFTBackbone=Stable Diffusion 1.52026.05 | 0.3389 | 0.2821 | 5.5981 | 0.583 | 0.2181 | — | — | |
| DSPOBackbone=Stable Diffusion 1.52026.05 | 0.3385 | 0.2819 | 5.5997 | 0.564 | 0.2178 | — | — | |
| DSPOBase Model=SD1.5, Samples per prompt=52026.05 | 0.3376 | 0.2813 | 5.658 | 0.5763 | 21.521 | — | — | |
| Diff.-DPOBase Model=SD1.5, Samples per prompt=52026.05 | 0.3373 | 0.2773 | 5.445 | 0.3539 | 21.497 | — | — | |
| MaPOBackbone=Stable Diffusion 1.52026.05 | 0.3366 | 0.2754 | 5.4754 | 0.3358 | 0.2152 | — | — | |
| Pre-TrainedBackbone=Stable Diffusion 1.52026.05 | 0.3343 | 0.2724 | 5.3466 | 0.0637 | 0.2144 | — | — | |
| SD1.5Base Model=SD1.5, Samples per prompt=52026.05 | 0.3322 | 0.2738 | 5.36 | 0.1653 | 21.243 | — | — | |
| FLUX-1.DevSpeed (secs/image)=12, Inference steps=502026.02 | 0.3197 | — | — | 1.192 | — | — | — | |
| DDiTSpeed (secs/image)=5.52026.02 | 0.3192 | — | — | 1.189 | — | — | — | |
| DDiT + TeacacheSpeed (secs/image)=3.4, delta=0.42026.02 | 0.3172 | — | — | 1.062 | — | — | — | |
| TaylorSeerSpeed (secs/image)=6, N=3, O=22026.02 | 0.3142 | — | — | 0.9813 | — | — | — | |
| FLUX-1.DevSpeed (secs/image)=6.72, Inference steps=282026.02 | 0.3118 | — | — | 1.115 | — | — | — | |
| FLUX-1.DevSpeed (secs/image)=3.6, Inference steps=152026.02 | 0.3072 | — | — | 0.9613 | — | — | — | |
| TeaCacheSpeed (secs/image)=6, delta=0.62026.02 | 0.3018 | — | — | 0.9701 | — | — | — | |
| TaylorSeerSpeed (secs/image)=3.5, N=6, O=12026.02 | 0.3004 | — | — | 0.9714 | — | — | — | |
| TeaCacheSpeed (secs/image)=5.33, delta=0.82026.02 | 0.2991 | — | — | 0.9699 | — | — | — | |
| CRAFTBase Model=SDXL2026.03 | — | 94.73 | 81.5 | 80.7 | — | — | — | |
| Diffusion-DPO2024.02 | — | 0.2815 | 5.7758 | 1.1495 | 22.2723 | 7.3698 | — | |
| DSPOBase Model=SDXL2026.03 | — | 81.8 | 57.84 | 73.47 | — | — | — | |
| SAILBase Model=SD1.5, Iteration=02026.02 | — | — | 5.28 | 0.1951 | 21.48 | — | 26.94 |