Question Answering on PIQA (Accuracy)
86.5AccuracyMashup Learning
Evaluation Results
| Method | Links | |
|---|---|---|
| Mashup LearningBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 86.5 | |
| BF16Models=Mixtral-8x7B, Quantization Strategy=Embedding-wise, Quantization Bit-width=BF162026.04 | 85.6 | |
| From scratchBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 85.2 | |
| Full-precisionModel=Mixtral 8x22B, Avg. bits/exp.=16 (FP), Memory (GB)=281.2, Evaluation protocol=zero-shot2026.04 | 85.12 | |
| Text-to-LoRA used as initBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 85.1 | |
| Full-precisionModel=Mixtral 8x7B, Avg. bits/exp.=16 (FP), Memory (GB)=96.8, Evaluation protocol=zero-shot2026.04 | 83.68 | |
| Gemma 2 9BModel=Gemma 2 9B2026.03 | 83 | |
| Focus (Gemma 2 9B)Backbone=Gemma 2 9B, K=4, dg=162026.03 | 83 | |
| FPModel=LLaDA-1.5, Quantization=Full Precision2026.06 | 82.97 | |
| FPModel=LLaDA-Instruct, Quantization=Full Precision2026.06 | 82.86 | |
| CodeQuantModels=Mixtral-8x7B, Quantization Strategy=Embedding-wise, Quantization Bit-width=A4W42026.04 | 82.7 | |
| FAIR-CalibModel=LLaDA-1.5, Quantization=W4A42026.06 | 82.7 | |
| Mistral 7BModel=Mistral 7B2026.03 | 82.6 | |
| Focus (Mistral 7B)Backbone=Mistral 7B, K=4, dg=162026.03 | 82.6 | |
| ICRBackbone=Qwen2.5-7B2025.09 | 82.6 | |
| LLaMA-2 70BModel=LLaMA-2 70B2026.03 | 82.4 | |
| Focus (LLaMA-2 70B)Backbone=LLaMA-2 70B, K=4, dg=162026.03 | 82.4 | |
| UniformModel=Mixtral 8x7B, Avg. bits/exp.=3, Memory (GB)=19.3, Evaluation protocol=zero-shot2026.04 | 82.32 | |
| FAIR-CalibModel=LLaDA-Instruct, Quantization=W4A42026.06 | 82.1 | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.75, Memory (GB)=17.7, Evaluation protocol=zero-shot2026.04 | 82.05 | |
| STLBase Model=LLaMA-3-8B2026.02 | 82 | |
| STL-augBase Model=LLaMA-3-8B2026.02 | 82 | |
| SCANSBase Model=LLaMA-3-8B2026.02 | 82 | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.75, Memory (GB)=17.7, Evaluation protocol=zero-shot2026.04 | 81.83 | |
| FlatQuantModel=LLaDA-Instruct, Quantization=W4A42026.06 | 81.83 | |
| BaselineBit width (b)=16.002025.05 | 81.8 | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.625, Memory (GB)=16.9, Evaluation protocol=zero-shot2026.04 | 81.56 | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.625, Memory (GB)=16.9, Evaluation protocol=zero-shot2026.04 | 81.45 | |
| UniformModel=Mixtral 8x22B, Avg. bits/exp.=3, Memory (GB)=57.5, Evaluation protocol=zero-shot2026.04 | 81.45 | |
| MoNEBackbone=Qwen1.5-MoE-A2.7B, Type=P25% Q4b, Storage=5.60GB, Evaluation Protocol=zero-shot2026.06 | 81.44 | |
| FlatQuantModel=LLaDA-1.5, Quantization=W4A42026.06 | 81.28 | |
| QuaRotModel=LLaDA-Instruct, Quantization=W4A42026.06 | 81.23 | |
| OursQBackbone=Qwen3-MoE-30B-A3B, Type=P25% Q4b, Storage=11.64GB, Evaluation Protocol=zero-shot2026.06 | 81.18 | |
| SurgicalBase Model=LLaMA-3-8B2026.02 | 81 | |
| OLMo-2 7BModel=OLMo-2 7B2026.03 | 81 | |
| Focus (OLMo-2 7B)Backbone=OLMo-2 7B, K=4, dg=162026.03 | 81 | |
| SARQC-GBS(Saliency)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.96 | |
| FP16Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 80.9 | |
| LLaMA-3 8BPruning Ratio=0%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 80.85 | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=16.1, Evaluation protocol=zero-shot2026.04 | 80.79 | |
| SARQC-GS(Identity)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.79 | |
| BaselineBackbone=Qwen1.5-MoE-A2.7B, Type=–, Storage=28.67GB, Evaluation Protocol=zero-shot2026.06 | 80.79 | |
| QuaRotModel=LLaDA-1.5, Quantization=W4A42026.06 | 80.74 | |
| Llama3 8BCompression Ratio=Baseline2025.09 | 80.7 | |
| Tensor RMS + CBit width (b)=3.002025.05 | 80.7 | |
| Llama3 8B2026.02 | 80.7 | |
| FP16Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.69 | |
| AWQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.69 | |
| SARQC-GS(Saliency)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.69 | |
| FP16Model=LLaMA-30B, Precision=W3A16, Zero-shot=true2026.05 | 80.69 | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=16.1, Evaluation protocol=zero-shot2026.04 | 80.63 | |
| Mashup Learning at initBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 80.6 | |
| SARQC-GBS(Identity)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.58 | |
| BF16Backbone=Llama-2-13B2026.05 | 80.52 | |
| BF16Models=Qwen3-30B-A3B, Quantization Strategy=Embedding-wise, Quantization Bit-width=BF162026.04 | 80.5 | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.375, Memory (GB)=15.3, Evaluation protocol=zero-shot2026.04 | 80.47 | |
| BaselineModel=Qwen3-MoE, Weight Bits=16, Zero-shot=true2026.04 | 80.47 | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.25, Memory (GB)=14.5, Evaluation protocol=zero-shot2026.04 | 80.41 | |
| LLaMA-2 13BModel=LLaMA-2 13B2026.03 | 80.4 | |
| Focus (LLaMA-2 13B)Backbone=LLaMA-2 13B, K=2, dg=162026.03 | 80.4 | |
| BF16Models=DeepSeek-V2-Lite, Quantization Strategy=Embedding-wise, Quantization Bit-width=BF162026.04 | 80.4 | |
| CodeQuantModel=DeepSeek-V2-Lite, Quantization Precision=A8W4, Quantization Granularity=Block-wise2026.04 | 80.4 | |
| BaselineBackbone=Mixtral, Weight Bitwidth=162026.04 | 80.36 | |
| GPTAQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.25 | |
| HessianModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=17.0, Evaluation protocol=zero-shot2026.04 | 80.21 | |
| Text-to-LoRABackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 80.2 | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.375, Memory (GB)=15.3, Evaluation protocol=zero-shot2026.04 | 80.2 | |
| Baseline (Dense)Sparsity=0.0%, Zero-shot=true2026.05 | 80.2 | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.25, Memory (GB)=14.5, Evaluation protocol=zero-shot2026.04 | 80.14 | |
| Router norm + Max varModel=Mixtral 8x22B, Avg. bits/exp.=2.5, Memory (GB)=46.7, Evaluation protocol=zero-shot2026.04 | 80.14 | |
| GPTQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 80.14 | |
| SARQC-GBS(Saliency)Model=LLaMA-30B, Precision=W3A16, Zero-shot=true2026.05 | 80.14 | |
| SARQC-GBS(Saliency)Model=LLaMA-7B, Precision=W4A16, Zero-shot=true2026.05 | 80.03 | |
| SARQC-GBS(Identity)Model=LLaMA-30B, Precision=W3A16, Zero-shot=true2026.05 | 80.03 | |
| STLBase Model=Qwen2.5-7B2026.02 | 80 | |
| STL-augBase Model=Qwen2.5-7B2026.02 | 80 | |
| SurgicalBase Model=Qwen2.5-7B2026.02 | 80 | |
| Qwen 2.5 7BModel=Qwen 2.5 7B2026.03 | 80 | |
| Focus (Qwen 2.5 7B)Backbone=Qwen 2.5 7B, K=4, dg=162026.03 | 80 | |
| FVBackbone=Qwen2.5-7B2025.09 | 80 | |
| SARQC-GBS(Saliency)Model=LLaMA-13B, Precision=W4A16, Zero-shot=true2026.05 | 79.92 | |
| WandaBackbone=Qwen1.5-MoE-A2.7B, Type=P25%, Storage=28.67GB, Evaluation Protocol=zero-shot2026.06 | 79.87 | |
| CodeQuantModel=DeepSeek-V2-Lite, Quantization Precision=A8W4, Quantization Granularity=Embedding-wise2026.04 | 79.8 | |
| BaselineModel=OLMoE, Weight Bits=16, Zero-shot=true2026.04 | 79.76 | |
| BaselineBackbone=Qwen1.5-MoE, Weight Bitwidth=162026.04 | 79.76 | |
| BF16Models=Phi-mini MoE-Instruct, Quantization Strategy=Embedding-wise, Quantization Bit-width=BF162026.04 | 79.7 | |
| CodeQuantModel=Qwen3-30B-A3B, Quantization Precision=A8W4, Quantization Granularity=Embedding-wise2026.04 | 79.7 | |
| PuzzleMoEBackbone=Qwen1.5-MoE-A2.7B, Type=M25%, Storage=22.40GB, Evaluation Protocol=zero-shot2026.06 | 79.7 | |
| FP16Model=LLaMA-13B, Precision=W4A16, Zero-shot=true2026.05 | 79.65 | |
| FP16Model=LLaMA-13B, Precision=W3A16, Zero-shot=true2026.05 | 79.65 | |
| CodeQuantModel=Phi-mini MoE-Instruct, Quantization Precision=A8W4, Quantization Granularity=Embedding-wise2026.04 | 79.6 | |
| BaselineBackbone=Qwen3-MoE-30B-A3B, Type=–, Storage=61.06GB, Evaluation Protocol=zero-shot2026.06 | 79.6 | |
| DuQuant++*Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 79.5 | |
| PMQModel=Mixtral 8x22B, Avg. bits/exp.=2.5, Memory (GB)=46.7, Evaluation protocol=zero-shot2026.04 | 79.49 | |
| CodeQuantModels=DeepSeek-V2-Lite, Quantization Strategy=Block-wise, Quantization Bit-width=A4W42026.04 | 79.4 | |
| PuzzleMoEBackbone=Qwen1.5-MoE-A2.7B, Type=M50%, Storage=16.17GB, Evaluation Protocol=zero-shot2026.06 | 79.4 | |
| GPTQModel=LLaMA-13B, Precision=W4A16, Zero-shot=true2026.05 | 79.38 | |
| SARQC-GBS(Identity)Model=LLaMA-13B, Precision=W4A16, Zero-shot=true2026.05 | 79.38 | |
| WandaBackbone=Qwen1.5-MoE-A2.7B, Type=P25% Q4b, Storage=7.10GB, Evaluation Protocol=zero-shot2026.06 | 79.38 | |
| OursBackbone=Qwen3-MoE-30B-A3B, Type=P50%, Storage=32.07GB, Evaluation Protocol=zero-shot2026.06 | 79.33 |