Inference Efficiency on LLaMA2-7B (12/128 tokens)
1.889LatencySWM (Ours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SWM (Ours)Pruned Rates=35%, Pruning Type=depth, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 1.889 | 67.77 | 8,725.9 | |
| SLEBPruned Rates=35%, Pruning Type=depth, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 1.938 | 66.048 | 8,725.9 | |
| Shortened-LLMPruned Rates=35%, Pruning Type=depth, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 2.084 | 61.433 | 8,725.85 | |
| SWM (Ours)Pruned Rates=20%, Pruning Type=depth, nparam=5.5B, input tokens=12, output tokens=128, batch size=12025.02 | 2.339 | 54.758 | 10,682.45 | |
| SLEBPruned Rates=20%, Pruning Type=depth, nparam=5.5B, input tokens=12, output tokens=128, batch size=12025.02 | 2.529 | 50.622 | 10,682.45 | |
| Shortened-LLMPruned Rates=20%, Pruning Type=depth, nparam=5.5B, input tokens=12, output tokens=128, batch size=12025.02 | 2.585 | 49.542 | 10,682.45 | |
| LLaMA2-7BPruned Rates=0%, nparam=6.7B, input tokens=12, output tokens=128, batch size=12025.02 | 2.729 | 46.905 | 13,020.25 | |
| FLAPPruned Rates=20%, Pruning Type=width, nparam=5.4B, input tokens=12, output tokens=128, batch size=12025.02 | 4.045 | 31.656 | 10,707.25 | |
| FLAPPruned Rates=35%, Pruning Type=width, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 4.127 | 31.051 | 8,855.95 | |
| Wanda-spPruned Rates=35%, Pruning Type=width, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 4.619 | 27.726 | 8,901 | |
| Wanda-spPruned Rates=20%, Pruning Type=width, nparam=5.5B, input tokens=12, output tokens=128, batch size=12025.02 | 4.628 | 27.663 | 10,676 | |
| LLM-PrunerPruned Rates=35%, Pruning Type=width, nparam=4.5B, input tokens=12, output tokens=128, batch size=12025.02 | 5.63 | 22.736 | 9,043.9 | |
| LLM-PrunerPruned Rates=20%, Pruning Type=width, nparam=5.5B, input tokens=12, output tokens=128, batch size=12025.02 | 5.655 | 22.635 | 10,951.5 |