Sparse Decoding (MHA) on Synthetic 128K Context (H100, FP16)
0.7FlashInfer Latency (ms)FlashInfer
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| FlashInferBatch size (B)=1, Attention Type=MHA, Hq=32, Hk=322026.05 | 0.7 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=4, Attention Type=MHA, Hq=32, Hk=322026.05 | 2.79 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=8, Attention Type=MHA, Hq=32, Hk=322026.05 | 5.61 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=16, Attention Type=MHA, Hq=32, Hk=322026.05 | 11.18 | — | — | — | — | — | — | |
| Sparse Decode (Double Sparsity)Batch size (B)=1, Attention Type=MHA, Hq=32, Hk=32, Indexer=Double Sparsity [48]2026.05 | — | 0.91 | 1.62 | 2.18 | 2.68 | 3.14 | 3.37 | |
| Sparse Decode (Double Sparsity)Batch size (B)=4, Attention Type=MHA, Hq=32, Hk=32, Indexer=Double Sparsity [48]2026.05 | — | 1.02 | 1.86 | 2.56 | 3.17 | 3.74 | 4 | |
| Sparse Decode (Double Sparsity)Batch size (B)=8, Attention Type=MHA, Hq=32, Hk=32, Indexer=Double Sparsity [48]2026.05 | — | 1.12 | 2 | 2.71 | 3.32 | 3.86 | 4.11 | |
| Sparse Decode (Double Sparsity)Batch size (B)=16, Attention Type=MHA, Hq=32, Hk=32, Indexer=Double Sparsity [48]2026.05 | — | 1.22 | 2.13 | 2.84 | 3.43 | 3.94 | 4.17 |