Large Language Model Evaluation on LLaMA-2 7B
5.47PPLbaseline
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| baseline#Bits=FP16, #Eff. (w-bits)=16, Training-free=✓, Latency overhead=81.4%2026.03 | 5.47 | 64.09 | 41.83 | |
| SERQ#Bits=W4A4, #Eff. (w-bits)=4.24, Training-free=✓, Latency overhead=18.7%2026.03 | 5.97 | 61.87 | 37.03 | |
| SpinQuant#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✗, Latency overhead=19.8%2026.03 | 6 | 61 | 34.8 | |
| QuaRot#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✓, Latency overhead=19.8%2026.03 | 6.15 | 59.53 | 33.58 | |
| SmoothQ(g128)#Bits=W4A4, #Eff. (w-bits)=4.13, Training-free=✓, Latency overhead=✗2026.03 | 7.49 | 57.15 | 30.4 | |
| SmoothQ#Bits=W4A4, #Eff. (w-bits)=∼4, Training-free=✓, Latency overhead=✗2026.03 | 15,100 | 35.44 | 26.04 |