Attention Mechanism Latency Benchmark on Synthetic sequences
0.03Latency (ms)Sparse Triton Kernel
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Sparse Triton KernelSequence Length=1024, Hardware=NVIDIA RTX 40902026.02 | 0.03 | 1.3 | 9.38 | |
| Sparse Triton KernelSequence Length=2048, Hardware=NVIDIA RTX 40902026.02 | 0.03 | 10.1 | 4.69 | |
| PyTorch SDPA (Flash Attention)Sequence Length=1024, Hardware=NVIDIA RTX 40902026.02 | 0.04 | — | — | |
| Sparse Triton KernelSequence Length=4096, Hardware=NVIDIA RTX 40902026.02 | 0.08 | 13.1 | 2.34 | |
| Sparse Triton KernelSequence Length=8192, Hardware=NVIDIA RTX 40902026.02 | 0.19 | 16.8 | 1.17 | |
| PyTorch SDPA (Flash Attention)Sequence Length=2048, Hardware=NVIDIA RTX 40902026.02 | 0.3 | — | — | |
| Sparse Triton KernelSequence Length=16384, Hardware=NVIDIA RTX 40902026.02 | 0.5 | 21.9 | 0.59 | |
| PyTorch SDPA (Flash Attention)Sequence Length=4096, Hardware=NVIDIA RTX 40902026.02 | 1 | — | — | |
| Sparse Triton KernelSequence Length=32768, Hardware=NVIDIA RTX 40902026.02 | 1.05 | 39.3 | 0.29 | |
| Sparse Triton KernelSequence Length=65536, Hardware=NVIDIA RTX 40902026.02 | 2.78 | 56.1 | 0.15 | |
| PyTorch SDPA (Flash Attention)Sequence Length=8192, Hardware=NVIDIA RTX 40902026.02 | 3.27 | — | — | |
| Sparse Triton KernelSequence Length=131072, Hardware=NVIDIA RTX 40902026.02 | 3.94 | 159.3 | 0.07 | |
| PyTorch SDPA (Flash Attention)Sequence Length=16384, Hardware=NVIDIA RTX 40902026.02 | 10.91 | — | — | |
| PyTorch SDPA (Flash Attention)Sequence Length=32768, Hardware=NVIDIA RTX 40902026.02 | 41.41 | — | — | |
| PyTorch SDPA (Flash Attention)Sequence Length=65536, Hardware=NVIDIA RTX 40902026.02 | 156.25 | — | — | |
| PyTorch SDPA (Flash Attention)Sequence Length=131072, Hardware=NVIDIA RTX 40902026.02 | 627.58 | — | — |