Quantization on OPT v1 (train)
0.08Processing Time (min)OPTQ
Evaluation Results
| Method | Links | |
|---|---|---|
| OPTQTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 0.08 | |
| OPTQTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 0.14 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 0.27 | |
| OPTQTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 0.29 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 0.61 | |
| OPTQTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 0.66 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 1.02 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 1.09 | |
| aespaTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 1.24 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 1.63 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 2.37 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 2.57 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 2.58 | |
| aespaTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 2.83 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 4.6 | |
| aespaTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 4.78 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 10.09 | |
| aespaTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 10.24 | |
| BRECQTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 10.71 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 16.2 | |
| BRECQTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 19.15 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 28.33 | |
| BRECQTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 108.2 |