Training Throughput on Synthetic 8192-token context (train)
229,957Training Throughput (tok/s)Full-sequence attention
Evaluation Results
| Method | Links | |
|---|---|---|
| Full-sequence attentionGPUs=8x A100, Global batch=64, Model size=350M, Context length=8192, Precision=BF16, Optimization=ZeRO-2, tiled linear cross-entropy, FlashAttention2026.07 | 229,957 | |
| Full-sequence DiffuMamba-HGPUs=8x A100, Global batch=64, Model size=350M, Context length=8192, Precision=BF16, Optimization=ZeRO-2, tiled linear cross-entropy, FlashAttention2026.07 | 223,825 | |
| Partially Reverse BDLM Mamba-HGPUs=8x A100, Global batch=64, Model size=350M, Context length=8192, Precision=BF16, Optimization=ZeRO-2, tiled linear cross-entropy, FlashAttention2026.07 | 166,849 | |
| BDLM attentionGPUs=8x A100, Global batch=64, Model size=350M, Context length=8192, Precision=BF16, Optimization=ZeRO-2, tiled linear cross-entropy, FlashAttention2026.07 | 151,976 |