ResearchBenchmarksLanguage Modeling Inference on Qwen2.5-7B 128K context lengthFollow18.4Decode Latency (ms/token)FastMKA16.80427.57738.3549.123Mar 21, 2026Evaluation ResultsMethodMethodLinksDecode Latency (ms/token)Speedup vs MLAFastMKABatch size=1, Precisio...Batch size=1, Precision=bf162026.0318.41.78MLABatch size=1, Precisio...Batch size=1, Precision=bf162026.0332.7—GQABatch size=1, Precisio...Batch size=1, Precision=bf162026.0349.8—MHABatch size=1, Precisio...Batch size=1, Precision=bf162026.0358.3—