Text-to-image generation on MoCA
2.607NIQEMoGen
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| MoGenInput=Text2026.01 | 2.607 | 0.342 | 92.47 | 4.857 | 4.492 | — | — | 65.28 | 67.75 | |
| MoGenInput=Text + Bounding Box2026.01 | 2.635 | 0.334 | 92.73 | 4.725 | 4.429 | — | 0.701 | 74.91 | 95.62 | |
| FLUXInput=Text2026.01 | 2.675 | 0.337 | 88.05 | 4.901 | 4.311 | — | — | 15.73 | 27.21 | |
| Omnigen2Input=Text2026.01 | 2.687 | 0.301 | 89.91 | 4.745 | 4.25 | — | — | 12.72 | 18.96 | |
| SDXLInput=Text2026.01 | 2.817 | 0.292 | 84.26 | 4.715 | 4.147 | — | — | 8.72 | 3.96 | |
| Emu2Input=Text + Bounding Box2026.01 | 2.921 | 0.307 | 85.29 | 4.515 | 4.106 | — | 0.472 | 25.49 | 0.57 | |
| Emu2Input=Text2026.01 | 2.991 | 0.297 | 83.77 | 4.491 | 4.089 | — | — | 6.72 | 2.31 | |
| Bounded-attentionInput=Text + Bounding Box2026.01 | 3.172 | 0.316 | 86.42 | 4.471 | 3.989 | — | 0.557 | 31.14 | 3.81 |