PrefixLM Attention on Qwen2.5 72B (q=64, k=8, 1k Context)
103.61PrefixLM Attention Throughput (TFLOPS)CuBridge
Evaluation Results
| Method | Links | |
|---|---|---|
| CuBridgeImplementation=CuBridge2026.05 | 103.61 | |
| FlexAttentionImplementation=PyTorch FlexAttention2026.05 | 99.82 | |
| Qimeng AttnImplementation=Qimeng Attention2026.05 | 72.37 | |
| TorchImplementation=Standard PyTorch2026.05 | 14.58 |