LLM Inference on 11B parameter model INT8
110Throughput (TPS)NVIDIA A100 PCIe 40GB
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| NVIDIA A100 PCIe 40GBClass=Datacenter, TDP (W)=250, BW (GB/s)=1555, Batch size=1, Quantization=INT82026.03 | 110 | 2.3 | |
| NVIDIA RTX 4090Class=Consumer, TDP (W)=450, BW (GB/s)=1008, Batch size=1, Quantization=INT82026.03 | 80 | 5.6 | |
| NVIDIA RTX 4090 (sys.)Class=Consumer, TDP (W)=∼600, BW (GB/s)=1008, Batch size=1, Quantization=INT82026.03 | 80 | 7.5 | |
| NVIDIA L4Class=Datacenter, TDP (W)=72, BW (GB/s)=300, Batch size=1, Quantization=INT82026.03 | 22 | 3.3 | |
| Apple M2 Pro (chip)Class=Consumer, TDP (W)=30, BW (GB/s)=200, Batch size=1, Quantization=INT82026.03 | 15 | 2 | |
| Apple M2 Pro (system)Class=Consumer, TDP (W)=∼40, BW (GB/s)=200, Batch size=1, Quantization=INT82026.03 | 15 | 2.7 |