LLM Inference on Meta-Llama-3-8B (128 token generation)
5,463Memory Usage (MB)AWQ (4-bit)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| AWQ (4-bit)Precision=4-bit2026.04 | 5,463 | 75.1 | 1,705 | 35.6 | 28.1 | 2.8 | 0.64 | |
| AWQ (3-bit)Precision=3-bit, Note=Identical to 4-bit2026.04 | 5,463 | 75.1 | 1,705 | 35.6 | 28.1 | 2.8 | 0.64 | |
| FASQ (Kernel)Effective Bit=3-bit, SZss=2, Ks=1282026.04 | 5,482 | 761.9 | 168 | 19.3 | 51.8 | 2.8 | 1.18 | |
| GPTQ (3-bit)Precision=3-bit2026.04 | 5,808 | 94 | 1,362 | 75.9 | 13.2 | 2.64 | 0.3 | |
| FASQ (Kernel)Effective Bit=4-bit, SZss=2, Ks=2562026.04 | 5,975 | 934.9 | 137 | 22.1 | 45.2 | 2.56 | 1.03 | |
| GPTQ (4-bit)Precision=4-bit2026.04 | 6,558 | 68.7 | 1,863 | 48.7 | 20.5 | 2.34 | 0.47 | |
| SmoothQuant W8A8Precision=W8A8, Note=W6A6 is identical2026.04 | 8,663 | 55.5 | 2,304 | 60.9 | 16.4 | 1.77 | 0.37 | |
| RTN (4-bit)Precision=4-bit2026.04 | 8,663 | 114.2 | 1,120 | 96.5 | 10.4 | 1.77 | 0.24 | |
| RTN (3-bit)Precision=3-bit2026.04 | 8,671 | 118.2 | 1,083 | 96.9 | 10.3 | 1.77 | 0.24 | |
| FP16 (Tensor Core)Precision=FP16, Implementation=Tensor Core2026.04 | 15,317 | 40.8 | 3,137 | 22.8 | 43.9 | 1 | 1 | |
| FASQ (reconstruct)Effective Bit=3-bit, SZss=2, Ks=1282026.04 | 15,474 | 302.1 | 424 | 285.7 | 3.5 | 0.99 | 0.08 | |
| FASQ (reconstruct)Effective Bit=4-bit, SZss=2, Ks=2562026.04 | 15,939 | 309.7 | 413 | 289.1 | 3.5 | 0.96 | 0.08 |