Text-to-Image Generation on DrawWaldoWorlds (Tier C)
56VQA AccuracyPeopleComposer
Evaluation Results
| Method | Links | |
|---|---|---|
| PeopleComposerInput Setting=Parsed structured inputs2026.05 | 56 | |
| Seg2AnyInput Setting=Parsed structured inputs2026.05 | 21 | |
| GrounDiTInput Setting=Parsed structured inputs2026.05 | 6 | |
| BoxDiffInput Setting=Parsed structured inputs2026.05 | 5 | |
| Layout GuidanceInput Setting=Parsed structured inputs2026.05 | 4 | |
| R&BInput Setting=Parsed structured inputs2026.05 | 3 | |
| InteractDiffusionInput Setting=Parsed structured inputs2026.05 | 3 |