Text-to-image on Standard text-to-image benchmarks
97.28CLIP ScoreBaseline Vanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Baseline VanillaVAE=Original VAE2025.11 | 97.28 | 29.44 | |
| Baseline HFVAE=Original VAE2025.11 | 97.27 | 30.08 | |
| StreamFlow (xformer+taesd+VFB)backend=Xformer, VAE=taesd2025.11 | 97.26 | 30.75 | |
| StreamFlow (xformer+Original VAE+VFB)backend=Xformer, VAE=Original VAE2025.11 | 96.74 | 31.82 | |
| StreamFlow (tensorrt+taesd+VFB)backend=TensorRT, VAE=taesd2025.11 | 96.25 | 31.28 | |
| StreamFlow (tensorrt+Original VAE+VFB)backend=TensorRT, VAE=Original VAE2025.11 | 96.08 | 31.16 |