Box-conditioned Image Generation on LVIS v1 (val)
44.6APUpper bound
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Upper boundInput=Real images, Evaluation Detector=ViTDet-L [33], Resolution=512x5122024.02 | 44.6 | 57.7 | 33.2 | 55 | 66.1 | 31.4 | 44.5 | 50.5 | |
| InstanceDiffusionInput Condition=Box inputs, Evaluation Protocol=Zero-shot2024.02 | 17.9 | 25.5 | 5.5 | 24.2 | 45 | 12.7 | 18.7 | 19.3 | |
| GLIGENInput Condition=Box inputs, Evaluation Protocol=Zero-shot, Reproduced=true2024.02 | 9.9 | 9.5 | 1.6 | 10.5 | 31.1 | 7.4 | 10 | 10.9 |