Text-to-Image Generation on DrawWaldoWorlds (Tier A)
84VQA AccuracyPeopleComposer
Evaluation Results
| Method | Links | |
|---|---|---|
| PeopleComposerInput Setting=Parsed structured inputs2026.05 | 84 | |
| Seg2AnyInput Setting=Parsed structured inputs2026.05 | 42 | |
| Layout GuidanceInput Setting=Parsed structured inputs2026.05 | 42 | |
| BoxDiffInput Setting=Parsed structured inputs2026.05 | 40 | |
| GrounDiTInput Setting=Parsed structured inputs2026.05 | 39 | |
| R&BInput Setting=Parsed structured inputs2026.05 | 29 | |
| InteractDiffusionInput Setting=Parsed structured inputs2026.05 | 25 |