Language Model Inference on vLLM benchmark (1024 prompts)
1.839Latency (s)Fisher-IntDim-E
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Fisher-IntDim-EHardware=NVIDIA H100, Framework=vLLM, Precision=bf16, Compression Ratio=50%, Max New Tokens=2562026.06 | 1.839 | 8,909 | 251 | 15.09 | 2.274 | |
| Qwen1.5-MoEHardware=NVIDIA H100, Framework=vLLM, Precision=bf16, Max New Tokens=2562026.06 | 2.233 | 7,336 | 187 | 26.67 | 2.689 |