Large Language Model Evaluation on LLaMA-2 13B
4.88Perplexitybaseline
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| baseline#Bits=FP16, #Eff. (w-bits)=16, Training-free=✓, Latency overhead=81.4%2026.03 | 4.88 | 66.53 | 52.04 | |
| SpinQuant#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✗, Latency overhead=19.8%2026.03 | 5.2 | 64.8 | 47.8 | |
| SERQ#Bits=W4A4, #Eff. (w-bits)=4.24, Training-free=✓, Latency overhead=18.7%2026.03 | 5.2 | 64.82 | 47.17 | |
| QuaRot#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✓, Latency overhead=19.8%2026.03 | 5.41 | 62.55 | 47.25 | |
| SmoothQ(g128)#Bits=W4A4, #Eff. (w-bits)=4.13, Training-free=✓, Latency overhead=✗2026.03 | 6.31 | 61.28 | 39.83 | |
| SmoothQ#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✓, Latency overhead=✗2026.03 | 13,200 | 34.53 | 23.85 |