Sparse Decoding (GQA) on Synthetic 128K context (test)
0.19FlashInfer Latency (ms)FlashInfer
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| FlashInferBatch size (B)=1, Attention Type=GQA, Hq=32, Hk=82026.05 | 0.19 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=4, Attention Type=GQA, Hq=32, Hk=82026.05 | 0.72 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=8, Attention Type=GQA, Hq=32, Hk=82026.05 | 1.5 | — | — | — | — | — | — | |
| FlashInferBatch size (B)=16, Attention Type=GQA, Hq=32, Hk=82026.05 | 3.08 | — | — | — | — | — | — | |
| Sparse Decode (Double Sparsity)Batch size (B)=1, Attention Type=GQA, Hq=32, Hk=8, Indexer=Double Sparsity [48]2026.05 | — | 0.28 | 0.56 | 0.83 | 1.12 | 1.46 | 1.65 | |
| Sparse Decode (Double Sparsity)Batch size (B)=4, Attention Type=GQA, Hq=32, Hk=8, Indexer=Double Sparsity [48]2026.05 | — | 0.32 | 0.67 | 1.06 | 1.52 | 2.11 | 2.45 | |
| Sparse Decode (Double Sparsity)Batch size (B)=8, Attention Type=GQA, Hq=32, Hk=8, Indexer=Double Sparsity [48]2026.05 | — | 0.36 | 0.75 | 1.18 | 1.68 | 2.3 | 2.66 | |
| Sparse Decode (Double Sparsity)Batch size (B)=16, Attention Type=GQA, Hq=32, Hk=8, Indexer=Double Sparsity [48]2026.05 | — | 0.41 | 0.85 | 1.31 | 1.82 | 2.46 | 2.81 |