Compositional Text-to-Image Generation on T2I-CompBench (BLIP-VQA and Human Preference)
0.3815BLIP-VQA Score (Color)SD 1.5 (w/ MACCO-CLIP text encoder)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SD 1.5 (w/ MACCO-CLIP text encoder)Base Model=SD 1.5, Text Encoder=MACCO-CLIP2026.06 | 0.3815 | 0.4236 | 0.3835 | -0.3295 | -0.384 | -0.2793 | |
| SD 1.5 (w/ vanilla CLIP text encoder)Base Model=SD 1.5, Text Encoder=vanilla CLIP2026.06 | 0.3651 | 0.4135 | 0.3721 | -0.4381 | -0.4349 | -0.3323 |