Throughput Benchmarking on vLLM (1k Input / 1k Output Tokens)
2.42Throughput (req/s)Fisher-MoE
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Fisher-MoEHardware=NVIDIA A100-80G, Input tokens=1,000, Output tokens=1,000, Serving framework=vLLM v0.8.4, Compression Ratio=50%2026.06 | 2.42 | 4,853.52 | 2.1 | 1.21 | |
| Qwen1.5-MoE-A2.7B-ChatHardware=NVIDIA A100-80G, Input tokens=1,000, Output tokens=1,000, Serving framework=vLLM v0.8.42026.06 | 2.01 | 4,010.27 | 1.74 | 1 | |
| Qwen1.5-7B-ChatHardware=NVIDIA A100-80G, Input tokens=1,000, Output tokens=1,000, Serving framework=vLLM v0.8.42026.06 | 1.15 | 2,298.89 | 1 | 0.57 |