Multimodal Inference on Qwen-Image (inference)
14.92Inference Latency (s)FastUSP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FastUSPGPUs=2, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x768, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 14.92 | 497.3 | 1.09 | |
| USP (para_attn)GPUs=2, Quantization=FP8 weights (optimum-quanto), Inference Steps=30, Resolution=1024x768, GPU=NVIDIA GeForce RTX 5090 (32GB), Interconnect=NVLink (~900 GB/s bidirectional)2026.02 | 16.33 | 544.4 | 1 |