ResearchBenchmarksLanguage Modeling Inference on Qwen2.5-7B (8K context length)Follow7.1Decode Latency (ms/token)FastMKA6.7129.33111.9514.569Mar 21, 2026Evaluation ResultsMethodMethodLinksDecode Latency (ms/token)Speedup vs MLAFastMKABatch size=1, Precisio...Batch size=1, Precision=bf162026.037.11.44MLABatch size=1, Precisio...Batch size=1, Precision=bf162026.0310.2—GQABatch size=1, Precisio...Batch size=1, Precision=bf162026.0314.8—MHABatch size=1, Precisio...Batch size=1, Precision=bf162026.0316.8—