Text-to-Image Generation on MARIO-Eval
34.7CLIPScoreTextDiffuser-2
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TextDiffuser-2short_name=TD-22024.11 | 34.7 | 33.8 | 56.2 | — | — | — | — | — | — | — | — | 7.17 | 4.97 | |
| GlyphControlsampling steps=50, classifier-free guidance=7.52023.11 | 34.56 | 50.82 | 32.56 | 64.07 | — | — | 31.37 | 33.33 | 21.35 | 15.38 | 29.67 | 18.18 | — | |
| TextDiffuser-2sampling steps=50, classifier-free guidance=7.52023.11 | 34.5 | 33.66 | 57.58 | 75.06 | 71.57 | 100 | 41.18 | 33.33 | 36.98 | 46.15 | 40.66 | 45.45 | — | |
| TextDiffusershort_name=TD2024.11 | 34.4 | 42 | 55.4 | — | — | — | — | — | — | — | — | 7.16 | 4.67 | |
| TextDiffusersampling steps=50, classifier-free guidance=7.52023.11 | 34.36 | 38.76 | 56.09 | 78.24 | 28.43 | 0 | 27.45 | 33.33 | 23.44 | 30.77 | 19.23 | 36.36 | — | |
| Type-RBase Model=Flux2024.11 | 33.1 | 43.1 | 62 | — | — | — | — | — | — | — | — | 8.55 | 7.67 | |
| Type-RBase Model=SD32024.11 | 33 | 45 | 48.6 | — | — | — | — | — | — | — | — | 8.46 | 7.3 | |
| SD-XLsampling steps=50, classifier-free guidance=7.52023.11 | 31.31 | 62.54 | 0.31 | 3.66 | — | — | — | — | 14.58 | 7.69 | 7.14 | 0 | — | |
| PixArt-αsampling steps=50, classifier-free guidance=7.52023.11 | 27.88 | 87.09 | 0.02 | 0.03 | — | — | — | — | 3.65 | 0 | 3.3 | 0 | — | |
| GlyphControlCapability=Image generation only2024.07 | 0.36 | 345 | — | — | — | — | — | — | — | — | — | — | — | |
| AnyTextCapability=Image generation only2024.07 | 0.36 | 352 | — | — | — | — | — | — | — | — | — | — | — | |
| TextHarmony-GenCapability=Text and image generation, Training mode=image generation/editing data only2024.07 | 0.36 | 330 | — | — | — | — | — | — | — | — | — | — | — | |
| TextDiffuser-2Capability=Image generation only2024.07 | 0.35 | 336 | — | — | — | — | — | — | — | — | — | — | — | |
| TextHarmonyCapability=Text and image generation, Slide-LoRA=true2024.07 | 0.35 | 342 | — | — | — | — | — | — | — | — | — | — | — | |
| SceneTextGen-7MPre-defined Text Layout=false2024.06 | 0.3455 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GlyphControl-10M*Pre-defined Text Layout=true, Training Dataset=LAION-Glyph2024.06 | 0.345 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TextDiffuser-10MPre-defined Text Layout=true2024.06 | 0.3436 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ControlNetPre-defined Text Layout=true2024.06 | 0.3424 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TextDiffuser-7MPre-defined Text Layout=false2024.06 | 0.3385 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TextHarmony*Capability=Text and image generation, Slide-LoRA=false2024.07 | 0.33 | 356 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepFloydPre-defined Text Layout=false2024.06 | 0.3267 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Latent Diffusion ModelPre-defined Text Layout=false2024.06 | 0.3015 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-InterleavedCapability=Text and image generation2024.07 | 0.29 | 412 | — | — | — | — | — | — | — | — | — | — | — | |
| SEED-LLaMA-14BCapability=Text and image generation2024.07 | 0.27 | 348 | — | — | — | — | — | — | — | — | — | — | — | |
| MiniGPT5Capability=Text and image generation2024.07 | 0.25 | 380 | — | — | — | — | — | — | — | — | — | — | — |