Boolean Question Answering on BoolQ
91.26AccuracyDirect Fine-tuning
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Direct Fine-tuningrank=1, training_steps=+0 steps2026.02 | 91.26 | — | — | |
| Direct Fine-tuningrank=1, training_steps=+200 steps2026.02 | 91 | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Min steps2026.02 | 90.92 | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+200 steps2026.02 | 90.81 | — | — | |
| Direct Fine-tuningrank=1, training_steps=+700 steps2026.02 | 90.79 | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+700 steps2026.02 | 90.56 | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Standard2026.02 | 90.29 | — | — | |
| Qwen3-30B-A3BPruning ratio=0%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 88.69 | — | — | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 87.8 | — | — | |
| FP16Base Model=Qwen3-8B2026.01 | 86.61 | — | — | |
| RSPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 86.61 | — | — | |
| Fullw_eff=∞2026.04 | 86.6 | — | — | |
| Stochasticw_eff=2562026.04 | 86.6 | — | — | |
| Qwen2-57B-A14BPruning ratio=0%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 86.45 | — | — | |
| HiF4+HiGPTQModel=Qwen2.5-14B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 86.27 | — | 0.73 | |
| MoBA (k=2)w_eff=2562026.04 | 86.1 | — | — | |
| HiF4Model=Qwen2.5-14B, A-W Quant Type=HiF42026.02 | 85.86 | — | 0.32 | |
| MC-SMoEPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85.84 | — | — | |
| NVFP4+PTSModel=Qwen2.5-14B, A-W Quant Type=NVFP4+PTS2026.02 | 85.72 | — | 0.18 | |
| BF16Model=Qwen2.5-14B, A-W Quant Type=BF162026.02 | 85.54 | — | — | |
| MoNEPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85.41 | — | — | |
| FAQBase Model=Qwen3-8B2026.01 | 85.29 | — | — | |
| FP16Base Model=Qwen3-4B2026.01 | 85.11 | — | — | |
| Stochasticw_eff=1282026.04 | 85.1 | — | — | |
| FP16Base Model=Qwen2.5-7B2026.01 | 84.71 | — | — | |
| SWAw_eff=2562026.04 | 84.7 | — | — | |
| UnifiedQA-largeparameters=770M, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 84.5 | — | — | |
| IgnoringBase model=LLaMA-2-13B2025.08 | 84.5 | — | — | |
| AWQBase Model=Qwen3-8B2026.01 | 84.22 | — | — | |
| ForgettingBase model=LLaMA-2-13B2025.08 | 84.13 | — | — | |
| NVFP4Model=Qwen2.5-14B, A-W Quant Type=NVFP42026.02 | 83.49 | — | -2.05 | |
| FAQBase Model=Qwen2.5-7B2026.01 | 83.3 | — | — | |
| MoBA (k=2)w_eff=1282026.04 | 82.6 | — | — | |
| Full Tokens (standard SFT)Base model=LLaMA-2-13B2025.08 | 82.24 | — | — | |
| HiF4+HiGPTQModel=Mistral-7B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 82.22 | — | 0.05 | |
| BaselineBit width (b)=16.002025.05 | 82.2 | — | — | |
| BF16Model=Mistral-7B, A-W Quant Type=BF162026.02 | 82.17 | — | — | |
| OriginalModel=LLaMA-3.1-8B, Ratio=1.0, Evaluation Protocol=0-shot2026.04 | 82.1 | — | — | |
| HiF4Model=Mistral-7B, A-W Quant Type=HiF42026.02 | 81.35 | — | -0.82 | |
| BaselineParam Ratio=1.00, Base Model=LLaMA-3-8b, Fine-tuning Protocol=None2025.12 | 81.3 | — | — | |
| BF16Model=Llama3-8B, A-W Quant Type=BF162026.02 | 81.16 | — | — | |
| MoNEPruning ratio=50%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 80.99 | — | — | |
| NVFP4+PTSModel=Mistral-7B, A-W Quant Type=NVFP4+PTS2026.02 | 80.86 | — | -1.31 | |
| UnifiedQA-baseparameters=220M, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 80.8 | — | — | |
| LLaMA2-7BNumber of Parameters=7B, Backbone=LLaMA2, Shots=02024.07 | 80.7 | — | — | |
| BaseBase model=LLaMA-2-13B2025.08 | 80.67 | — | — | |
| AWQBase Model=Qwen2.5-7B2026.01 | 80.49 | — | — | |
| HiF4+HiGPTQModel=Llama3-8B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 80.4 | — | -0.76 | |
| MoonlightPruning ratio=0%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 80.4 | — | — | |
| Llama 2 7Bshot=zero-shot, normalization=none2024.02 | 80.2 | — | — | |
| AWQBase Model=Qwen3-4B2026.01 | 80 | — | — | |
| MC-SMoEPruning ratio=50%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 79.91 | — | — | |
| Deepseek-V2-LitePruning ratio=0%, Base Model=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 79.88 | — | — | |
| HiF4Model=Llama3-8B, A-W Quant Type=HiF42026.02 | 79.85 | — | -1.31 | |
| SFTVocab size=64k, Number of shots=4-shot2025.12 | 79.8 | — | — | |
| ESAMBackbone=Llama3-3B-Instruct, Training Mode=Fine-tuning2026.02 | 79.79 | — | — | |
| RTNBase Model=Qwen3-8B2026.01 | 79.48 | — | — | |
| CD-MoE-SR#L=9, ACT.=2.8B, MEM.=72.5%, Speedup=1.26x, SFT Protocol=w/ lightweight SFT2024.11 | 79.4 | — | — | |
| SL-SAMBackbone=Llama3-3B-Instruct, Training Mode=Fine-tuning2026.02 | 79.39 | — | — | |
| Qwen3-1.7B#Params=1.7B/1.7B, #Tokens=36T, Zero-shot=true2026.02 | 79.3 | — | — | |
| ALM*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 79.3 | — | — | |
| RSTBackbone=Llama3-3B-Instruct, Training Mode=Fine-tuning2026.02 | 79.24 | — | — | |
| AdaSAMBackbone=Llama3-3B-Instruct, Training Mode=Fine-tuning2026.02 | 79.2 | — | — | |
| AdamWBackbone=Llama3-3B-Instruct, Training Mode=Fine-tuning2026.02 | 79.14 | — | — | |
| FKL*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.9 | — | — | |
| DSKD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.9 | — | — | |
| FKL*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.9 | — | — | |
| FKLVocab size=32k, Number of shots=4-shot2025.12 | 78.8 | — | — | |
| SFTVocab size=16k, Number of shots=4-shot2025.12 | 78.8 | — | — | |
| OLTQAConfiguration=Full2023.05 | 78.78 | — | — | |
| ALM*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.6 | — | — | |
| DSKD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.6 | — | — | |
| ALMVocab size=16k, Number of shots=4-shot2025.12 | 78.6 | — | — | |
| ALM*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.6 | — | — | |
| FKL*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.5 | — | — | |
| FAQBase Model=Qwen3-4B2026.01 | 78.41 | — | — | |
| ULDVocab size=32k, Number of shots=4-shot2025.12 | 78.4 | — | — | |
| NVFP4Model=Llama3-8B, A-W Quant Type=NVFP42026.02 | 78.35 | — | -2.81 | |
| SFTVocab size=32k, Number of shots=4-shot2025.12 | 78.3 | — | — | |
| ALMVocab size=32k, Number of shots=4-shot2025.12 | 78.3 | — | — | |
| ALMVocab size=64k, Number of shots=4-shot2025.12 | 78.2 | — | — | |
| ULD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 78.2 | — | — | |
| FKLVocab size=16k, Number of shots=4-shot2025.12 | 78.1 | — | — | |
| OriginalModel=LLaMA-2-7B, Ratio=1.0, Evaluation Protocol=0-shot2026.04 | 77.9 | — | — | |
| BaselineParam Ratio=1.00, Base Model=LLaMA-2-7b, Fine-tuning Protocol=None2025.12 | 77.8 | — | — | |
| BF16Model=Llama2-7B, A-W Quant Type=BF162026.02 | 77.74 | — | — | |
| RTNBase Model=Qwen2.5-7B2026.01 | 77.71 | — | — | |
| G-LionShots=3, Optimizer=Global Lion, Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 77.5 | — | — | |
| FKLVocab size=64k, Number of shots=4-shot2025.12 | 77.5 | — | — | |
| NVFP4+PTSModel=Llama3-8B, A-W Quant Type=NVFP4+PTS2026.02 | 77.46 | — | -3.7 | |
| SPES-9B#Params=3.1B/9B, #Tokens=400B, Zero-shot=true, initialized from=pretrained dense model2026.02 | 77.3 | — | — | |
| LLRCParam Ratio=0.90, Base Model=LLaMA-2-7b, Fine-tuning Protocol=None2025.12 | 77.3 | — | — | |
| G-AdamWShots=3, Optimizer=Global AdamW, Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 77.23 | — | — | |
| DSKDVocab size=64k, Number of shots=4-shot2025.12 | 77.2 | — | — | |
| DSKDVocab size=32k, Number of shots=4-shot2025.12 | 77.2 | — | — | |
| Distributed Lion-MaVoShots=3, Optimizer=D-Lion (MaVo), Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 77.14 | — | — | |
| Stochasticw_eff=642026.04 | 77.1 | — | — | |
| OLTQAKnowledge Distillation=Static MKD2023.05 | 77.03 | — | — | |
| Distributed Lion-AvgShots=3, Optimizer=D-Lion (Avg), Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 76.9 | — | — | |
| HiF4Model=Llama2-7B, A-W Quant Type=HiF42026.02 | 76.73 | — | -1.01 |