ResearchBenchmarksLanguage Modeling Inference on Qwen2.5-7B (4K Context Length)Follow6.2Decode Latency (ms/token)FastMKA5.888.0410.212.36Mar 21, 2026Evaluation ResultsMethodMethodLinksDecode Latency (ms/token)Speedup vs MLAFastMKABatch size=1, Precisio...Batch size=1, Precision=bf162026.036.21.4MLABatch size=1, Precisio...Batch size=1, Precision=bf162026.038.7—GQABatch size=1, Precisio...Batch size=1, Precision=bf162026.0312.4—MHABatch size=1, Precisio...Batch size=1, Precision=bf162026.0314.2—