LLM Inference Serving Performance on 1024 Input / 32 Output Context
1.48TPOT Speedup vs DeepGEMMRaMP
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| RaMPRequest rate (r)=4, Model=OLMoE-1B-7B, Hardware=H200, Precision=FP8, Prompts per workload=802026.04 | 1.48 | 1.34 | 1.16 | 1.12 | 1.19 | 1.06 | |
| RaMPRequest rate (r)=2, Model=OLMoE-1B-7B, Hardware=H200, Precision=FP8, Prompts per workload=802026.04 | 1.43 | 1.34 | 1.15 | 1.04 | 1.18 | 1.08 | |
| RaMPRequest rate (r)=8, Model=OLMoE-1B-7B, Hardware=H200, Precision=FP8, Prompts per workload=802026.04 | 1.24 | 1.29 | 1.11 | 0.95 | 1.13 | 1.02 |