Text-to-Image Generation on DrawWaldoWorlds (Tier B) (VQA Accuracy)
72VQA AccuracyPeopleComposer
Evaluation Results
| Method | Links | |
|---|---|---|
| PeopleComposerInput Setting=Parsed structured inputs2026.05 | 72 | |
| Seg2AnyInput Setting=Parsed structured inputs2026.05 | 36 | |
| BoxDiffInput Setting=Parsed structured inputs2026.05 | 15 | |
| Layout GuidanceInput Setting=Parsed structured inputs2026.05 | 14 | |
| GrounDiTInput Setting=Parsed structured inputs2026.05 | 13 | |
| InteractDiffusionInput Setting=Parsed structured inputs2026.05 | 8 | |
| R&BInput Setting=Parsed structured inputs2026.05 | 7 |