Text-to-Image Generation on InterActing Multi-subject Interaction scenario sampled
3.8Human Likert ScoreDetailScribe
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| DetailScribeBackbone=Stable Diffusion 3.52025.04 | 3.8 | 4.6 | 1.326 | 0.907 | 34.3 | |
| DALL·E 32025.04 | 3.775 | 4.7 | 1.111 | 0.813 | 28.6 | |
| SD + Inf ScaleBackbone=Stable Diffusion 3.5, Strategy=Inference Scaling2025.04 | 3.375 | 4.4 | 1.149 | 0.903 | 30.2 | |
| SD + GPT RewriteBackbone=Stable Diffusion 3.5, Prompt Strategy=GPT Rewrite2025.04 | 3.275 | 3.9 | 1.101 | 0.89 | 29.2 | |
| SDBackbone=Stable Diffusion 3.52025.04 | 3.225 | 3.8 | 1.171 | 0.894 | 30.6 | |
| SD + GPT RefineBackbone=Stable Diffusion 3.5, Prompt Strategy=GPT Refine2025.04 | 3.175 | 3.6 | 1.022 | 0.858 | 29.2 |