Inference Throughput on 700M random-initialized models
1,935Throughput (Seq Len 256)BDLM attention
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BDLM attentionHardware=1x A100-80GB, Precision=BF16, Batch size=1, Block size=256, Optimized cached denoising=true, CUDA graphs usage=graph-captured block steps2026.07 | 1,935 | 1,887 | 1,955 | 1,927 | 1,728 | 1,663 | 1,331 | 508 | 263 | 142 | 91 | 74 | |
| Partially Reverse BDLM Mamba-HHardware=1x A100-80GB, Precision=BF16, Batch size=1, Block size=256, Optimized cached denoising=true, CUDA graphs usage=graph-captured block steps2026.07 | 1,146 | 1,167 | 1,082 | 1,158 | 1,019 | 1,281 | 1,234 | 941 | 694 | 450 | 364 | 278 | |
| Full-sequence attentionHardware=1x A100-80GB, Precision=BF16, Batch size=1, CUDA graphs usage=fixed-shape denoising2026.07 | 1,035 | 915 | 1,421 | 1,062 | 761 | 351 | 122 | 36 | — | — | — | — | |
| Full-sequence DiffuMamba-HHardware=1x A100-80GB, Precision=BF16, Batch size=1, CUDA graphs usage=fixed-shape denoising2026.07 | 470 | 541 | 498 | 461 | 387 | 429 | 240 | 101 | 35 | — | — | — |