Text-to-Image Generation on FLUX.1 12B inference 1.0 (dev)
4.67Inference Time (s)FastUSP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FastUSPGPUs=8, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 4.67 | 155.7 | 1.12 | |
| USP (para_attn)GPUs=8, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 5.22 | 174 | 1 | |
| FastUSPGPUs=4, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 6.25 | 208.3 | 1.12 | |
| USP (para_attn)GPUs=4, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 7 | 233.3 | 1 | |
| FastUSPGPUs=2, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 7.46 | 248.7 | 1.16 | |
| USP (para_attn)GPUs=2, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x1024, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 8.64 | 288 | 1 |