Training Throughput Analysis on Qwen 7B 2.5
1,847Training Throughput (tokens/s)FastMKA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FastMKASequence Length=4K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,847 | 3.94 | |
| FastMKASequence Length=8K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,732 | 4.05 | |
| FastMKASequence Length=16K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,614 | 4.17 | |
| FastMKASequence Length=32K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,453 | 4.25 | |
| FastMKASequence Length=64K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,287 | 4.48 | |
| FastMKASequence Length=128K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 1,032 | 4.87 | |
| FastMKASequence Length=256K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 742 | 5.01 | |
| MLASequence Length=4K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 468 | — | |
| MLASequence Length=8K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 428 | — | |
| GQASequence Length=4K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 424 | — | |
| MLASequence Length=16K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 387 | — | |
| GQASequence Length=8K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 382 | — | |
| MHASequence Length=4K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 347 | — | |
| MLASequence Length=32K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 342 | — | |
| GQASequence Length=16K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 341 | — | |
| GQASequence Length=32K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 296 | — | |
| MHASequence Length=8K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 289 | — | |
| MLASequence Length=64K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 287 | — | |
| GQASequence Length=64K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 243 | — | |
| MHASequence Length=16K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 231 | — | |
| MLASequence Length=128K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 212 | — | |
| MHASequence Length=32K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 184 | — | |
| GQASequence Length=128K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 178 | — | |
| MLASequence Length=256K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 148 | — | |
| MHASequence Length=64K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 142 | — | |
| GQASequence Length=256K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 124 | — | |
| MHASequence Length=128K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 98 | — | |
| MHASequence Length=256K, Batch Size=2, Precision=bf16, Attention Implementation=FlashAttention-22026.03 | 68 | — |