Quantization on OPT
4.8Processing Time (s)OPTQ
Evaluation Results
| Method | Links | |
|---|---|---|
| OPTQTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 4.8 | |
| OPTQTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 4.8 | |
| OPTQTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 8.4 | |
| OPTQTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 8.4 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 16.2 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 16.2 | |
| OPTQTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 17.4 | |
| OPTQTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 17.4 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 36.6 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 36.6 | |
| OPTQTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 39.6 | |
| OPTQTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 39.6 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 61.2 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 61.2 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 65.4 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 65.4 | |
| aespaTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 74.4 | |
| aespaTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 74.4 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 97.8 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 97.8 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 142.2 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 154.2 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 154.2 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 154.8 | |
| Z-FOLDTarget=layer-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 154.8 | |
| aespaTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 169.8 | |
| aespaTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 169.8 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 175.8 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 276 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 276 | |
| aespaTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 286.8 | |
| aespaTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 286.8 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 591 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 591 | |
| aespaTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 614.4 | |
| aespaTarget=attention-wise reconstruction, Model Size=6.7B, Bit-width=INT22024.02 | 614.4 | |
| BRECQTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 642.6 | |
| BRECQTarget=attention-wise reconstruction, Model Size=1.3B, Bit-width=INT22024.02 | 642.6 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 972 | |
| OmniQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 972 | |
| BRECQTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 1,149 | |
| BRECQTarget=attention-wise reconstruction, Model Size=2.7B, Bit-width=INT22024.02 | 1,149 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 1,699.8 | |
| AffineQuantTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 1,699.8 | |
| BRECQTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 6,492 | |
| BRECQTarget=attention-wise reconstruction, Model Size=125M, Bit-width=INT22024.02 | 6,492 |