LLM Decoding on Llama-2 70B (Throughput)
36.1Throughput (tokens/s)AWQ
Evaluation Results
| Method | Links | |
|---|---|---|
| AWQQuantization Configuration=2BIT, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 36.1 | |
| LIFTQUANTQuantization Configuration=LQ-32/16, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 31.3 | |
| LIFTQUANTQuantization Configuration=LQ-24/10, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 25.7 | |
| QTIPQuantization Configuration=2BIT, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 24.5 | |
| LIFTQUANTQuantization Configuration=LQ-24/8, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 20.8 | |
| QTIPQuantization Configuration=3BIT, Hardware=GTX 4090D (48GB), Context Length=512, Batch size=12026.06 | 17.6 |