Text-to-Image Generation on Visual Genome (train test)
38.61Inception ScoreVG
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VGDescription Source=Real Visual Genome images2024.05 | 38.61 | 35.32 | 0 | — | |
| GT Synth.Description Source=GT training scene descriptions, Rendering Pipeline=Stable Diffusion2024.05 | 24.44 | 0 | 35.32 | — | |
| GCE frameworkPrompt Enrichment Strategy=Graph-based enrichment (Ours), Rendering Pipeline=Stable Diffusion2024.05 | 19.37 | 34.89 | 53.23 | 13.74 | |
| SimplePrompt Enrichment Strategy=Direct rendering from input description, Rendering Pipeline=Stable Diffusion2024.05 | 17.03 | 37.85 | 61.27 | — | |
| GPT ObjectPrompt Enrichment Strategy=Enriched ChatGPT Object, Rendering Pipeline=Stable Diffusion2024.05 | 15.69 | 52.74 | 73.48 | — | |
| GPT DirectPrompt Enrichment Strategy=Enriched ChatGPT Direct, Rendering Pipeline=Stable Diffusion2024.05 | 15.44 | 42.15 | 64.69 | — | |
| GPT ScenePrompt Enrichment Strategy=Enriched ChatGPT Scene, Rendering Pipeline=Stable Diffusion2024.05 | 14.9 | 47.74 | 68.06 | — |