Quantization on LLAMA
0.25Processing Time (hr)OPTQ
Evaluation Results
| Method | Links | |
|---|---|---|
| OPTQTarget=layer-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 0.25 | |
| OPTQTarget=layer-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 0.25 | |
| OPTQTarget=layer-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 0.45 | |
| OPTQTarget=layer-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 0.45 | |
| OPTQTarget=layer-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 1.08 | |
| OPTQTarget=layer-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 1.08 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 1.13 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 1.13 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 2.37 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 2.37 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 2.48 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 2.48 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 4.2 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 4.2 | |
| aespaTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 6.84 | |
| aespaTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 6.84 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 9.84 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 9.84 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 10.09 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=7B, Bit-width=INT22024.02 | 10.09 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 10.51 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 10.51 | |
| aespaTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 15.89 | |
| aespaTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 15.89 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 18.76 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=13B, Bit-width=INT22024.02 | 18.76 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 47.84 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 47.84 | |
| aespaTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 53.69 | |
| aespaTarget=attention-wise reconstruction, Model Size=30B, Bit-width=INT22024.02 | 53.69 |