Attention Operator Throughput on Llama 405B (128 Q-heads/8 KV-heads/128 Head-dimension) 3.1
615.39TFLOPSCuBridge
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CuBridgeAttention Variant=ReLU Attention2026.05 | 615.39 | 3.12 | — | — | — | — | |
| CuBridgeAttention Variant=Sigmoid Attention2026.05 | 578 | 3.65 | — | — | — | — | |
| FlashInferAttention Variant=Causal Mask, Implementation Mode=Native2026.05 | 551.62 | — | — | — | — | — | |
| CuBridgeAttention Variant=Relative Position2026.05 | 549.43 | 3.36 | — | — | — | — | |
| CuBridgeAttention Variant=Causal Mask2026.05 | 546.18 | 0.99 | — | — | — | — | |
| CuBridgeAttention Variant=PrefixLM2026.05 | 515.44 | 3 | — | — | — | — | |
| CuBridgeAttention Variant=Combo2026.05 | 501.43 | 5.59 | — | — | — | — | |
| CuBridgeAttention Variant=Softcap2026.05 | 466.85 | 1.22 | — | — | — | — | |
| CuBridgeAttention Variant=ALiBi2026.05 | 442.71 | 1.07 | — | — | — | — | |
| FlashInferAttention Variant=ALiBi, Implementation Mode=Native2026.05 | 412.12 | — | — | — | — | — | |
| CuBridgeAttention Variant=Share Question Mask2026.05 | 411.71 | 4.15 | — | — | — | — | |
| CuBridgeAttention Variant=Causal Blockwise Mask2026.05 | 388.72 | 3.11 | — | — | — | — | |
| FlashInferAttention Variant=Softcap, Implementation Mode=Native2026.05 | 383.9 | — | — | — | — | — | |
| CuBridgeAttention Variant=Sliding Window2026.05 | 282.6 | 1.02 | — | — | — | — | |
| FlashInferAttention Variant=Sliding Window, Implementation Mode=Native2026.05 | 276.12 | — | — | — | — | — | |
| flash-attn v2Sequence Length=16k, Hardware=A100 GPU, Head Dimension=1282025.06 | 225.3 | — | — | — | — | — | |
| CuBridgeAttention Variant=Global Sliding Window2026.05 | 225.25 | 2.73 | — | — | — | — | |
| flash-attn v2Sequence Length=8k, Hardware=A100 GPU, Head Dimension=1282025.06 | 221.1 | — | — | — | — | — | |
| cuDNNSequence Length=16k, Hardware=A100 GPU, Head Dimension=1282025.06 | 211.2 | — | — | — | — | — | |
| flash-attn v2Sequence Length=4k, Hardware=A100 GPU, Head Dimension=1282025.06 | 209.1 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=16k, Hardware=A100 GPU, Head Dimension=1282025.06 | 206.9 | — | — | — | — | — | |
| cuDNNSequence Length=8k, Hardware=A100 GPU, Head Dimension=1282025.06 | 206.4 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=8k, Hardware=A100 GPU, Head Dimension=1282025.06 | 204.4 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=4k, Hardware=A100 GPU, Head Dimension=1282025.06 | 200.3 | — | — | — | — | — | |
| FlashInferAttention Variant=ReLU Attention, Implementation Mode=JIT2026.05 | 197.06 | — | — | — | — | — | |
| cuDNNSequence Length=4k, Hardware=A100 GPU, Head Dimension=1282025.06 | 196.7 | — | — | — | — | — | |
| flash-attn v2Sequence Length=2k, Hardware=A100 GPU, Head Dimension=1282025.06 | 192.2 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=2k, Hardware=A100 GPU, Head Dimension=1282025.06 | 187.5 | — | — | — | — | — | |
| cuDNNSequence Length=2k, Hardware=A100 GPU, Head Dimension=1282025.06 | 183.4 | — | — | — | — | — | |
| FlashInferAttention Variant=Causal Mask, Implementation Mode=JIT2026.05 | 176.54 | — | — | — | — | — | |
| FlexAttentionSequence Length=16k, Hardware=A100 GPU, Head Dimension=1282025.06 | 175.3 | — | — | — | — | — | |
| FlashInferAttention Variant=PrefixLM, Implementation Mode=JIT2026.05 | 171.56 | — | — | — | — | — | |
| FlexAttentionSequence Length=8k, Hardware=A100 GPU, Head Dimension=1282025.06 | 169.7 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=1k, Hardware=A100 GPU, Head Dimension=1282025.06 | 168.4 | — | — | — | — | — | |
| flash-attn v2Sequence Length=1k, Hardware=A100 GPU, Head Dimension=1282025.06 | 164.1 | — | — | — | — | — | |
| FlashInferAttention Variant=Relative Position, Implementation Mode=JIT2026.05 | 163.61 | — | — | — | — | — | |
| FlashInferAttention Variant=Sigmoid Attention, Implementation Mode=JIT2026.05 | 158.25 | — | — | — | — | — | |
| FlexAttentionSequence Length=4k, Hardware=A100 GPU, Head Dimension=1282025.06 | 157.7 | — | — | — | — | — | |
| cuDNNSequence Length=1k, Hardware=A100 GPU, Head Dimension=1282025.06 | 157.2 | — | — | — | — | — | |
| QiMeng-AttentionSequence Length=512, Hardware=A100 GPU, Head Dimension=1282025.06 | 146.5 | — | — | — | — | — | |
| FlexAttentionSequence Length=2k, Hardware=A100 GPU, Head Dimension=1282025.06 | 144.4 | — | — | — | — | — | |
| flash-attn v2Sequence Length=512, Hardware=A100 GPU, Head Dimension=1282025.06 | 129.1 | — | — | — | — | — | |
| FlashInferAttention Variant=Causal Blockwise Mask, Implementation Mode=BSR2026.05 | 125 | — | — | — | — | — | |
| cuDNNSequence Length=512, Hardware=A100 GPU, Head Dimension=1282025.06 | 121.9 | — | — | — | — | — | |
| FlexAttentionSequence Length=1k, Hardware=A100 GPU, Head Dimension=1282025.06 | 117.2 | — | — | — | — | — | |
| FlashInferAttention Variant=PrefixLM, Implementation Mode=BSR2026.05 | 111.1 | — | — | — | — | — | |
| FlashInferAttention Variant=Causal Mask, Implementation Mode=BSR2026.05 | 110.95 | — | — | — | — | — | |
| FlashInferAttention Variant=Share Question Mask, Implementation Mode=BSR2026.05 | 99.28 | — | — | — | — | — | |
| FlexAttentionSequence Length=512, Hardware=A100 GPU, Head Dimension=1282025.06 | 93.2 | — | — | — | — | — | |
| FlashInferAttention Variant=Softcap, Implementation Mode=JIT2026.05 | 91.2 | — | — | — | — | — | |
| FlashInferAttention Variant=Combo, Implementation Mode=JIT2026.05 | 89.76 | — | — | — | — | — | |
| FlashInferAttention Variant=ALiBi, Implementation Mode=JIT2026.05 | 85.37 | — | — | — | — | — | |
| FlashInferAttention Variant=Global Sliding Window, Implementation Mode=BSR2026.05 | 82.61 | — | — | — | — | — | |
| FlashInferAttention Variant=Sliding Window, Implementation Mode=BSR2026.05 | 78.92 | — | — | — | — | — | |
| FlashInferAttention Variant=Global Sliding Window, Implementation Mode=JIT2026.05 | 68.28 | — | — | — | — | — | |
| FlashInferAttention Variant=Sliding Window, Implementation Mode=JIT2026.05 | 38.27 | — | — | — | — | — | |
| FlashInferAttention Variant=Causal Blockwise Mask, Implementation Mode=JIT2026.05 | 35.67 | — | — | — | — | — | |
| FlashInferAttention Variant=Share Question Mask, Implementation Mode=JIT2026.05 | 11.89 | — | — | — | — | — | |
| DeepSeek-V3Sequence Length=1k, Hardware=A100 GPU, Head Dimension=1282025.06 | 11.2 | — | — | — | — | — | |
| DeepSeek-V3Sequence Length=512, Hardware=A100 GPU, Head Dimension=1282025.06 | 11.1 | — | — | — | — | — | |
| DeepSeek-V3Sequence Length=4k, Hardware=A100 GPU, Head Dimension=1282025.06 | 10.2 | — | — | — | — | — | |
| DeepSeek-V3Sequence Length=2k, Hardware=A100 GPU, Head Dimension=1282025.06 | 8.9 | — | — | — | — | — | |
| CuBridgeHardware=NVIDIA H100, Heads (q, k)=128, 82026.05 | — | — | 383.51 | 548.49 | 532.16 | 597.61 | |
| CuBridgeHardware=NVIDIA H1002026.05 | — | — | 471.97 | 595.68 | 540.23 | 589.83 | |
| FlexAttentionHardware=NVIDIA H100, Heads (q, k)=128, 82026.05 | — | — | 280.66 | 335.76 | 377.96 | 400.05 | |
| FlexAttentionHardware=NVIDIA H1002026.05 | — | — | 334.1 | 380 | 407.89 | 429.1 | |
| Qimeng AttnHardware=NVIDIA H100, Heads (q, k)=128, 82026.05 | — | — | 124.68 | 173.7 | 214.14 | 238.76 | |
| Qimeng AttnHardware=NVIDIA H1002026.05 | — | — | 193.11 | 238.75 | 265.39 | 281.31 | |
| TorchHardware=NVIDIA H100, Heads (q, k)=128, 82026.05 | — | — | 29.03 | 20.06 | 25.67 | 29.66 | |
| TorchHardware=NVIDIA H1002026.05 | — | — | 33.45 | 26.5 | 32.52 | — |