ResearchBenchmarksEfficiency Evaluation on GPT-2 124MFollow1Inference Speed (x)FP32 Baseline0.950.97511.025Jan 5, 2026Evaluation ResultsMethodMethodLinksInference Speed (x)Model Size (MB)FP32 BaselinePrecision=float32, Qua...Precision=float32, Quantization=None2026.011—