PrefixLM Attention Performance on Llama2-7B (8k Context, 32/32 Heads)
163.7TFLOPS (PrefixLM Attention)CuBridge
Evaluation Results
| Method | Links | |
|---|---|---|
| CuBridgeImplementation=CuBridge2026.05 | 163.7 | |
| FlexAttentionImplementation=PyTorch FlexAttention2026.05 | 142.77 | |
| Qimeng AttnImplementation=Qimeng Attention2026.05 | 122.44 | |
| TorchImplementation=Standard PyTorch2026.05 | 14.56 |