Language Modeling on WikiText-2 (Perplexity)
3.84PerplexityVanilla
Evaluation Results
| Method | Links | |
|---|---|---|
| VanillaBackbone=Mixtral-8x7B-v0.1, Sparsity=0%2025.11 | 3.84 | |
| PuzzleMoEBackbone=Mixtral-8x7B-v0.1, Sparsity=25%2025.11 | 4.1 | |
| PuzzleMoEBackbone=Mixtral-8x7B-v0.1, Sparsity=50%2025.11 | 4.36 | |
| D2Backbone=Mixtral-8x7B-v0.1, Sparsity=20%2025.11 | 4.65 | |
| FP16Bits (W-A-KV)=16-16-16, Model=Q2.5-32B2025.06 | 4.67 | |
| NAEEBackbone=Mixtral-8x7B-v0.1, Sparsity=25%2025.11 | 5.01 | |
| FlatQuantBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.05 | |
| FPTQuantBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.12 | |
| Sub-MoEBackbone=Mixtral-8x7B-v0.1, Sparsity=25%2025.11 | 5.16 | |
| D2Backbone=Mixtral-8x7B-v0.1, Sparsity=40%2025.11 | 5.28 | |
| QuaRot-optBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.28 | |
| SpinQuantBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.28 | |
| OSTQuantBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.28 | |
| HC-SMoEBackbone=Mixtral-8x7B-v0.1, Sparsity=25%2025.11 | 5.31 | |
| DenseSparsity=0%, Hardware=A100, Batch Size=1, Sequence Length=5122026.07 | 5.33 | |
| FP16Bits (W-A-KV)=16-16-16, Model=L2-7B2025.06 | 5.47 | |
| BaselineRatio=0.0, Backbone=LLaMA-7B, Fine-tuning=false2026.07 | 5.68 | |
| RTN-optBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 5.72 | |
| FP16Bits (W-A-KV)=16-16-16, Model=L3-8B2025.06 | 5.75 | |
| FPTQuantBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 5.85 | |
| WandaBackbone=Mixtral-8x7B-v0.1, Sparsity=2:42025.11 | 5.89 | |
| SpinQuantBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 5.97 | |
| FlatQuantBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 6.2 | |
| QuaRot-optBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 6.22 | |
| FPTQuantBits (W-A-KV)=4-4-4, Model=Q2.5-32B2025.06 | 6.24 | |
| FPTQuantBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 6.27 | |
| FP16Bits (W-A-KV)=16-16-16, Model=M-8B it2025.06 | 6.45 | |
| FlatQuantBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 6.46 | |
| NAEEBackbone=Mixtral-8x7B-v0.1, Sparsity=50%2025.11 | 6.49 | |
| OSTQuantBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 6.49 | |
| VanillaBackbone=Deepseek-MoE-16b, Sparsity=0%2025.11 | 6.51 | |
| SpinQuantBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 6.54 | |
| OSTQuantBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 6.56 | |
| FlatQuantBits (W-A-KV)=4-4-4, Model=Q2.5-32B2025.06 | 6.57 | |
| PuzzleMoEBackbone=Deepseek-MoE-16b, Sparsity=25%2025.11 | 6.68 | |
| FlatQuantBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 6.69 | |
| FPTQuantBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 6.72 | |
| QuaRot-optBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 6.76 | |
| OSTQuantBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 6.82 | |
| RTNBits (W-A-KV)=4-8-8, Model=Q2.5-32B2025.06 | 6.83 | |
| D2Backbone=Deepseek-MoE-16b, Sparsity=20%2025.11 | 6.84 | |
| SpinQuantBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 6.86 | |
| PuzzleMoEBackbone=Deepseek-MoE-16b, Sparsity=50%2025.11 | 6.88 | |
| Sub-MoEBackbone=Mixtral-8x7B-v0.1, Sparsity=50%2025.11 | 6.97 | |
| QuaRot-optBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 7.04 | |
| RTN-optBits (W-A-KV)=4-8-8, Model=L2-7B2025.06 | 7.11 | |
| VanillaBackbone=Qwen1.5-MoE-A2.7B, Sparsity=0%2025.11 | 7.22 | |
| RTN-optBits (W-A-KV)=4-8-8, Model=L3-8B2025.06 | 7.32 | |
| PuzzleMoEBackbone=Qwen1.5-MoE-A2.7B, Sparsity=25%2025.11 | 7.37 | |
| LACE-SVDRatio=0.2, Backbone=LLaMA-7B, Fine-tuning=false2026.07 | 7.39 | |
| FlatQuantModel=Q2.5-7B it, Quantization Setting=(i) Linear+KV, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.47 | |
| FlatQuantModel=Q2.5-7B it, Quantization Setting=(ii) +BMM, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.51 | |
| QuaRot-optBits (W-A-KV)=4-4-4, Model=Q2.5-32B2025.06 | 7.51 | |
| PuzzleMoEBackbone=Qwen1.5-MoE-A2.7B, Sparsity=50%2025.11 | 7.55 | |
| SpinQuantBits (W-A-KV)=4-4-4, Model=Q2.5-32B2025.06 | 7.59 | |
| LACE-SVDMemory Budget=10 GB, Backbone=LLaMA-7B2026.07 | 7.6 | |
| FPTQuantModel=Q2.5-7B it, Quantization Setting=(i) Linear+KV, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.61 | |
| HC-SMoEBackbone=Mixtral-8x7B-v0.1, Sparsity=50%2025.11 | 7.65 | |
| SpinquantModel=Q2.5-7B it, Quantization Setting=(i) Linear+KV, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.66 | |
| FPTQuantModel=Q2.5-7B it, Quantization Setting=(ii) +BMM, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.74 | |
| SpinquantModel=Q2.5-7B it, Quantization Setting=(ii) +BMM, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 7.87 | |
| D2Backbone=Deepseek-MoE-16b, Sparsity=40%2025.11 | 7.93 | |
| SVD-LLMRatio=0.2, Backbone=LLaMA-7B, Fine-tuning=false2026.07 | 7.94 | |
| QuaRot-optBits (W-A-KV)=4-4-4, Model=M-8B it2025.06 | 8.34 | |
| OSTQuantBits (W-A-KV)=4-4-4, Model=Q2.5-32B2025.06 | 8.39 | |
| FPTQuantModel=Q2.5-7B it, Quantization Setting=(iii) All except residual, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 8.44 | |
| FlatQuantBits (W-A-KV)=4-4-4, Model=M-8B it2025.06 | 8.44 | |
| WandaBackbone=Deepseek-MoE-16b, Sparsity=2:42025.11 | 8.46 | |
| Sub-MoEBackbone=Deepseek-MoE-16b, Sparsity=25%2025.11 | 8.48 | |
| FPTQuantBits (W-A-KV)=4-4-4, Model=M-8B it2025.06 | 8.49 | |
| Dobi-SVDRatio=0.2, Backbone=LLaMA-7B, Fine-tuning=false2026.07 | 8.54 | |
| FP16BIT=162026.07 | 8.55 | |
| SpinQuantBits (W-A-KV)=4-4-4, Model=M-8B it2025.06 | 8.69 | |
| VanillaBackbone=Qwen3-MoE-30B-A3B, Sparsity=0%2025.11 | 8.7 | |
| SliceGPTMemory Budget=10 GB, Backbone=LLaMA-7B2026.07 | 8.78 | |
| WandaBackbone=Qwen1.5-MoE-A2.7B, Sparsity=2:42025.11 | 8.81 | |
| LACE-SVDMemory Budget=9 GB, Backbone=LLaMA-7B2026.07 | 8.85 | |
| GSRQBIT=22026.07 | 8.88 | |
| VQLLMBIT=22026.07 | 9.05 | |
| PuzzleMoEBackbone=Qwen3-MoE-30B-A3B, Sparsity=25%2025.11 | 9.08 | |
| SpinquantModel=Q2.5-7B it, Quantization Setting=(iii) All except residual, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 9.23 | |
| FlatQuantModel=Q2.5-7B it, Quantization Setting=(iii) All except residual, Quantization Precision (W-A-KV)=W4A8KV82025.06 | 9.24 | |
| BlockPrunerMemory Budget=10 GB, Backbone=LLaMA-7B2026.07 | 9.4 | |
| PuzzleMoEBackbone=Qwen3-MoE-30B-A3B, Sparsity=50%2025.11 | 9.5 | |
| FlatQuantModel=L3-8B, Quantization Setting=(i) Linear+KV, Quantization Precision (W-A-KV)=W4A4KV42025.06 | 9.55 | |
| FlatQuantBits (W-A-KV)=4-4-4, Model=L3-8B2025.06 | 9.55 | |
| FP16 BaselineBackbone=LLaMA3-1B, Weight bit-width=16, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 9.6 | |
| OSTQuantBits (W-A-KV)=4-4-4, Model=L3-8B2025.06 | 9.66 | |
| FPTQuantModel=L3-8B, Quantization Setting=(i) Linear+KV, Quantization Precision (W-A-KV)=W4A4KV42025.06 | 9.74 | |
| FPTQuantBits (W-A-KV)=4-4-4, Model=L3-8B2025.06 | 9.74 | |
| LLM-PrunerMemory Budget=10 GB, Backbone=LLaMA-7B2026.07 | 9.88 | |
| OSTQuantBits (W-A-KV)=4-4-4, Model=M-8B it2025.06 | 10 | |
| RTN-optBits (W-A-KV)=4-8-8, Model=M-8B it2025.06 | 10.13 | |
| GSRQBIT=12026.07 | 10.22 | |
| FP16Bits (W-A-KV)=16-16-16, Model=L3.2-3B it2025.06 | 10.48 | |
| FPTQuantBits (W-A-KV)=4-8-8, Model=L3.2-3B it2025.06 | 10.65 | |
| FlatQuantBits (W-A-KV)=4-8-8, Model=L3.2-3B it2025.06 | 10.67 | |
| QuaRot-optBits (W-A-KV)=4-8-8, Model=L3.2-3B it2025.06 | 10.89 | |
| D2Backbone=Qwen1.5-MoE-A2.7B, Sparsity=20%2025.11 | 10.91 | |
| PALSSparsity=50%, Hardware=A100, Batch Size=1, Sequence Length=5122026.07 | 10.96 |