Causal Blockwise Mask Attention on Llama2-7b (q=32, k=32) (1k)
35.12TFLOPSCuBridge
Evaluation Results
| Method | Links | |
|---|---|---|
| CuBridgeImplementation=CuBridge2026.05 | 35.12 | |
| FlexAttentionImplementation=PyTorch FlexAttention2026.05 | 31.77 | |
| Qimeng AttnImplementation=Qimeng Attention2026.05 | 16.91 | |
| TorchImplementation=Standard PyTorch2026.05 | 4.85 |