Multi-Domain Knowledge on MMLU
80.75MMLU Multi-Domain Knowledge AccFP16
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FP16Model=Qwen3-32B-Instruct, #Bits(W)=16, #Bits(A)=162026.05 | 80.75 | — | |
| FP16Model=Qwen3-32B-Instruct, Weight bits=16, Activation bits=162026.06 | 80.75 | — | |
| BF16Model=Qwen2.5-14B, A-W Quant Type=BF162026.02 | 80.17 | — | |
| HiF4+HiGPTQModel=Qwen2.5-14B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 78.65 | -1.52 | |
| NVFP4+PTSModel=Qwen2.5-14B, A-W Quant Type=NVFP4+PTS2026.02 | 78.63 | -1.54 | |
| HiF4Model=Qwen2.5-14B, A-W Quant Type=HiF42026.02 | 78.62 | -1.55 | |
| NVFP4Model=Qwen2.5-14B, A-W Quant Type=NVFP42026.02 | 76.53 | -3.64 | |
| TWLAModel=Qwen3-32B-Instruct, Weight bits=1.58, Activation bits=162026.06 | 75.59 | — | |
| GPTQModel=Qwen3-32B-Instruct, Weight bits=3, Activation bits=162026.06 | 73.97 | — | |
| TWLAModel=Qwen3-32B-Instruct, Weight bits=1.58, Activation bits=4MP2026.06 | 70.21 | — | |
| BWLAModel=Qwen3-32B-Instruct, #Bits(W)=1.15, #Bits(A)=162026.05 | 67.74 | — | |
| BWLAModel=Qwen3-32B-Instruct, #Bits(W)=1.15, #Bits(A)=62026.05 | 67.48 | — | |
| BF16Model=Llama3-8B, A-W Quant Type=BF162026.02 | 66.55 | — | |
| GPTQModel=Qwen3-32B-Instruct, #Bits(W)=3, #Bits(A)=162026.05 | 63.75 | — | |
| HiF4+HiGPTQModel=Llama3-8B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 63.55 | -3 | |
| HiF4Model=Llama3-8B, A-W Quant Type=HiF42026.02 | 63.37 | -3.18 | |
| BF16Model=Mistral-7B, A-W Quant Type=BF162026.02 | 63.3 | — | |
| NVFP4+PTSModel=Llama3-8B, A-W Quant Type=NVFP4+PTS2026.02 | 62.81 | -3.74 | |
| ARB-LLMModel=Qwen3-32B-Instruct, #Bits(W)=1.06, #Bits(A)=162026.05 | 62.56 | — | |
| NVFP4Model=Llama3-8B, A-W Quant Type=NVFP42026.02 | 61.64 | -4.91 | |
| HiF4Model=Mistral-7B, A-W Quant Type=HiF42026.02 | 61.63 | -1.67 | |
| HiF4+HiGPTQModel=Mistral-7B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 61.53 | -1.77 | |
| NVFP4+PTSModel=Mistral-7B, A-W Quant Type=NVFP4+PTS2026.02 | 61.15 | -2.15 | |
| T-SPINIteration=Iter42026.01 | 58.55 | — | |
| T-SPINIteration=Iter32026.01 | 58.51 | — | |
| SFTMethod=Supervised Fine-Tuning2026.01 | 58.37 | — | |
| SPINIteration=Iter42026.01 | 58.05 | — | |
| SPINIteration=Iter02026.01 | 57.97 | — | |
| Mistral-7BConfiguration=Base2026.01 | 57.86 | — | |
| SPINIteration=Iter32026.01 | 57.8 | — | |
| T-SPINIteration=Iter02026.01 | 57.74 | — | |
| SPINIteration=Iter22026.01 | 57.74 | — | |
| T-SPINIteration=Iter42026.01 | 57.68 | — | |
| T-SPINIteration=Iter32026.01 | 57.67 | — | |
| SPINIteration=Iter12026.01 | 57.63 | — | |
| T-SPINIteration=Iter22026.01 | 57.61 | — | |
| SPINIteration=Iter22026.01 | 57.49 | — | |
| T-SPINIteration=Iter12026.01 | 57.48 | — | |
| SFT2026.01 | 57.29 | — | |
| Zephyr-7B2026.01 | 56.9 | — | |
| T-SPINIteration=Iter12026.01 | 56.89 | — | |
| T-SPINIteration=Iter22026.01 | 56.89 | — | |
| SPINIteration=Iter12026.01 | 56.86 | — | |
| T-SPINIteration=Iter02026.01 | 56.42 | — | |
| SPINIteration=Iter02026.01 | 56.25 | — | |
| SPINIteration=Iter32026.01 | 55.88 | — | |
| SliM-LLMModel=Qwen3-32B-Instruct, Weight bits=2MP, Activation bits=162026.06 | 55.38 | — | |
| SPINIteration=Iter42026.01 | 53.59 | — | |
| BiLLMModel=Qwen3-32B-Instruct, #Bits(W)=1.06, #Bits(A)=162026.05 | 49.04 | — | |
| PB-LLMModel=Qwen3-32B-Instruct, Weight bits=1.7, Activation bits=162026.06 | 48.21 | — | |
| BF16Model=Llama2-7B, A-W Quant Type=BF162026.02 | 46.52 | — | |
| HiF4Model=Llama2-7B, A-W Quant Type=HiF42026.02 | 44.54 | -1.98 | |
| HiF4+HiGPTQModel=Llama2-7B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 43.89 | -2.63 | |
| NVFP4+PTSModel=Llama2-7B, A-W Quant Type=NVFP4+PTS2026.02 | 43.43 | -3.09 | |
| NVFP4Model=Llama2-7B, A-W Quant Type=NVFP42026.02 | 41.51 | -5.01 | |
| GPTQ + QIGModel=LLaMA-2-7B, Bit-width=3-bit2026.03 | 32.01 | — | |
| GPTQModel=LLaMA-2-7B, Bit-width=3-bit2026.03 | 30.05 | — | |
| GPTQModel=Qwen3-32B-Instruct, Weight bits=2, Activation bits=162026.06 | 27.1 | — | |
| NVFP4Model=Mistral-7B, A-W Quant Type=NVFP42026.02 | 26.79 | -36.51 | |
| BiLLMModel=Qwen3-32B-Instruct, #Bits(W)=1.06, #Bits(A)=62026.05 | 25.82 | — | |
| GPTQModel=Qwen3-32B-Instruct, #Bits(W)=3, #Bits(A)=62026.05 | 25.36 | — | |
| GPTQModel=Qwen3-32B-Instruct, #Bits(W)=2, #Bits(A)=62026.05 | 24.75 | — | |
| GPTQModel=Qwen3-32B-Instruct, #Bits(W)=2, #Bits(A)=162026.05 | 24.58 | — | |
| ARB-LLMModel=Qwen3-32B-Instruct, #Bits(W)=1.06, #Bits(A)=62026.05 | 24.33 | — | |
| SliM-LLMModel=Qwen3-32B-Instruct, Weight bits=2MP, Activation bits=42026.06 | 23.71 | — | |
| GPTQModel=Qwen3-32B-Instruct, Weight bits=3, Activation bits=42026.06 | 23.28 | — | |
| GPTQModel=Qwen3-32B-Instruct, Weight bits=2, Activation bits=42026.06 | 23.2 | — | |
| PB-LLMModel=Qwen3-32B-Instruct, Weight bits=1.7, Activation bits=42026.06 | 23.12 | — |