ResearchBenchmarksEnd-to-end LLM Inference Serving on ShareGPTFollow1.5TPOT Speedup vs DeepGEMMRaMP1.41681.43841.461.4816Apr 28, 2026Evaluation ResultsMethodMethodLinksTPOT Speedup vs DeepGEMMTPOT Speedup vs TritonTPOT Speedup vs FlashInferTTFT Speedup vs DeepGEMMTTFT Speedup vs TritonTTFT Speedup vs FlashInferRaMPRequest rate (r)=2, Mo...Request rate (r)=2, Model=OLMoE-1B-7B, Hardware=H200, Precision=FP8, Prompts per workload=802026.041.51.31.161.441.211.09RaMPRequest rate (r)=4, Mo...Request rate (r)=4, Model=OLMoE-1B-7B, Hardware=H200, Precision=FP8, Prompts per workload=802026.041.421.211.091.351.151.06