Large Language Model Evaluation on Qwen3-0.6B (test)
47.83Average PerformanceEDGERAZOR
Evaluation Results
| Method | Links | |
|---|---|---|
| EDGERAZORQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 47.83 | |
| BF16Quantization Configuration=W16-A16-KV16, Weight bit-width=16, Activation bit-width=16, KV cache bit-width=162026.04 | 47.35 | |
| AQLMQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 46.48 | |
| AutoRoundQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 45.75 | |
| AWQQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 44.65 | |
| GPTAQQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 44.49 | |
| EDGERAZORQuantization Configuration=W2.79-A16-KV16, Weight bit-width=2.79, Activation bit-width=16, KV cache bit-width=162026.04 | 44.17 | |
| GPTQQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 43.71 | |
| VPTQQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 41.69 | |
| EDGERAZORQuantization Configuration=W1.88-A16-KV16, Weight bit-width=1.88, Activation bit-width=16, KV cache bit-width=162026.04 | 41.6 | |
| Q-PaletteQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 40.97 | |
| AutoRoundQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 40.96 | |
| AQLMQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 39.85 | |
| EDGERAZORQuantization Configuration=W1.58-A16-KV16, Weight bit-width=1.58, Activation bit-width=16, KV cache bit-width=162026.04 | 39.77 | |
| Q-PaletteQuantization Configuration=W3.25-A16-KV16, Weight bit-width=3.25, Activation bit-width=16, KV cache bit-width=162026.04 | 37.55 | |
| VPTQQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 37.46 | |
| OmniQuantQuantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 36.6 | |
| AQLMQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 36.51 | |
| QTIPQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 35.94 | |
| GPTAQQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 35.61 | |
| QuIP#Quantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 35.42 | |
| AWQQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 35.37 | |
| OmniQuantQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 34.57 | |
| GPTQQuantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 34.53 | |
| Slim-LLM+Quantization Configuration=W3-A16-KV16, Weight bit-width=3, Activation bit-width=16, KV cache bit-width=162026.04 | 33.95 | |
| AutoRoundQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 31.8 | |
| VPTQQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 31.42 | |
| AWQQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 31.02 | |
| Q-PaletteQuantization Configuration=W1.75-A16-KV16, Weight bit-width=1.75, Activation bit-width=16, KV cache bit-width=162026.04 | 30.81 | |
| ARB-LLMQuantization Configuration=W1.00-A16-KV16, Weight bit-width=1.00, Activation bit-width=16, KV cache bit-width=162026.04 | 30.77 | |
| OmniQuantQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 30.7 | |
| Q-PaletteQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 30.66 | |
| Slim-LLM+Quantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 30.54 | |
| QuIP#Quantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 30.07 | |
| GPTQQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 30 | |
| BiLLMQuantization Configuration=W1.06-A16-KV16, Weight bit-width=1.06, Activation bit-width=16, KV cache bit-width=162026.04 | 29.98 | |
| QuIP#Quantization Configuration=W4-A16-KV16, Weight bit-width=4, Activation bit-width=16, KV cache bit-width=162026.04 | 29.9 | |
| GPTAQQuantization Configuration=W2-A16-KV16, Weight bit-width=2, Activation bit-width=16, KV cache bit-width=162026.04 | 29.8 |