Inference Efficiency on Model Profiling
15.9Total GEMM Latency (ms)ESPACE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ESPACEBackbone=GPT3-1.3B, Compression Ratio=47%, # of Weights=6.42 x 10^82024.10 | 15.9 | 31.7 | |
| ESPACEBackbone=GPT3-1.3B, Compression Ratio=20%, # of Weights=9.71 x 10^82024.10 | 20.6 | 36.1 | |
| BaselineBackbone=GPT3-1.3B, Compression Ratio=0%, # of Weights=1.21 x 10^92024.10 | 24.2 | 39.8 | |
| ESPACEBackbone=GPT3-8B, Compression Ratio=50%, # of Weights=4.08 x 10^92024.10 | 76.8 | 122 | |
| ESPACEBackbone=GPT3-8B, Compression Ratio=21%, # of Weights=6.33 x 10^92024.10 | 110 | 155 | |
| ESPACEBackbone=Llama2-7B, Compression Ratio=50%, # of Weights=3.24 x 10^92024.10 | 113 | 266 | |
| BaselineBackbone=GPT3-8B, Compression Ratio=0%, # of Weights=8.05 x 10^92024.10 | 136 | 186 | |
| ESPACEBackbone=Llama2-7B, Compression Ratio=21%, # of Weights=5.11 x 10^92024.10 | 169 | 322 | |
| ESPACEBackbone=GPT3-22B, Compression Ratio=55%, # of Weights=9.74 x 10^92024.10 | 181 | 261 | |
| BaselineBackbone=Llama2-7B, Compression Ratio=0%, # of Weights=6.48 x 10^92024.10 | 210 | 368 | |
| RetrainedBackbone=Llama2-7B, Compression Ratio=0%, # of Weights=6.48 x 10^92024.10 | 210 | 368 | |
| ESPACEBackbone=Nemotron4-15B, Compression Ratio=50%, # of Weights=6.25 x 10^92024.10 | 223 | 545 | |
| ESPACEBackbone=GPT3-22B, Compression Ratio=40%, # of Weights=1.30 x 10^102024.10 | 229 | 313 | |
| ESPACEBackbone=Llama2-13B, Compression Ratio=50%, # of Weights=6.34 x 10^92024.10 | 259 | 447 | |
| ESPACEBackbone=Nemotron4-15B, Compression Ratio=25%, # of Weights=9.54 x 10^92024.10 | 324 | 655 | |
| ESPACEBackbone=Llama2-13B, Compression Ratio=20%, # of Weights=1.01 x 10^102024.10 | 336 | 562 | |
| BaselineBackbone=GPT3-22B, Compression Ratio=0%, # of Weights=2.17 x 10^102024.10 | 354 | 457 | |
| BaselineBackbone=Llama2-13B, Compression Ratio=0%, # of Weights=1.27 x 10^102024.10 | 406 | 643 | |
| BaselineBackbone=Nemotron4-15B, Compression Ratio=0%, # of Weights=1.25 x 10^102024.10 | 414 | 741 |