Question Answering on BoolQ (Accuracy)
90.03AccuracyShortGPT
Evaluation Results
| Method | Links | |
|---|---|---|
| ShortGPTModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 90.03 | |
| ShortGPTModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.85 | |
| IT-PrunModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.66 | |
| IT-PrunModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.64 | |
| SLEBModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.6 | |
| IT-PrunModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.18 | |
| ShortGPTModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.14 | |
| SafeMERGEBackbone=Qwen2.5-7B2026.06 | 89.1 | |
| BEAMModel=Qwen3-30B-A3B, Beta (β)=0.01, Avg. K=4.232026.05 | 88.07 | |
| Fine-tunedBackbone=Qwen2.5-7B2026.06 | 87.98 | |
| SafeGeneBackbone=Qwen2.5-7B2026.06 | 87.86 | |
| SafeLoRABackbone=Qwen2.5-7B2026.06 | 86.97 | |
| SaLoRABackbone=Qwen2.5-7B2026.06 | 86.9 | |
| SLEBModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 86.76 | |
| Qwen3-30B-A3BModel=Qwen3-30B-A3B, Target Top-K (K)=8, Avg. K=8.002026.05 | 86.76 | |
| MKAModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 86.15 | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct, Sparsity=Dense2026.06 | 85.81 | |
| FP16*#Bits=16-16, Evaluation Protocol=original paper result2026.05 | 85.23 | |
| FP16#Bits=16-162026.05 | 85.05 | |
| GEMQ#Bits=3-162026.05 | 85.02 | |
| DenseBackbone=LLaMA-1, Model Parameters=65B, Sparsity=0%, Evaluation Protocol=zero-shot, Quantization=8-bit2025.11 | 84.83 | |
| Qwen3-4B-InstructModel=Qwen3-4B-Instruct, Sparsity=Dense2026.06 | 84.83 | |
| AutoPruneBackbone=LLaMA-1, Model Parameters=65B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot, Quantization=8-bit2025.11 | 84.79 | |
| WandaBackbone=LLaMA-1, Model Parameters=65B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot, Quantization=8-bit2025.11 | 84.7 | |
| SparseGPTBackbone=LLaMA-1, Model Parameters=65B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot, Quantization=8-bit2025.11 | 84.6 | |
| SUBFITModel=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 84.43 | |
| DeepSeek-7B-chatModel=DeepSeek-7B-chat, Sparsity=Dense2026.06 | 83.98 | |
| Llama-3.1-8B-InstructModel=Llama-3.1-8B-Instruct, Sparsity=Dense2026.06 | 83.88 | |
| AutoPruneParams=LLaMA-2 70B, Sparsity=50% unstructured, Evaluation protocol=zero-shot, Quantization=8-bit2025.11 | 83.6 | |
| SparseGPTParams=LLaMA-2 70B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 83.55 | |
| DenseParams=LLaMA-2 70B, Sparsity=0%, Evaluation protocol=zero-shot2025.11 | 83.4 | |
| FP16Backbone=LLaMA-3-8B, Evaluation Protocol=zero-shot2026.06 | 83.36 | |
| SUBFITModel=DeepSeek-7B-chat, Sparsity=25%2026.06 | 83.09 | |
| FP16Backbone=Qwen-3-8B, Evaluation Protocol=zero-shot2026.06 | 83.03 | |
| MoEQuant*#Bits=3-16, Evaluation Protocol=original paper result2026.05 | 82.81 | |
| DenseBackbone=LLaMA-1, Model Parameters=30B, Sparsity=0%, Evaluation Protocol=zero-shot2025.11 | 82.69 | |
| WandaParams=LLaMA-2 70B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 82.5 | |
| SparseGPTBackbone=LLaMA-1, Model Parameters=30B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 82.32 | |
| AutoPruneBackbone=LLaMA-1, Model Parameters=30B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 82.11 | |
| WandaBackbone=LLaMA-1, Model Parameters=30B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 81.9 | |
| WandaParams=LLaMA-2 13B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 81.84 | |
| SparseGPTParams=LLaMA-2 13B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 81.44 | |
| DenseModel=LLaMA-2-7B, Sparsity Pattern=Dense (None), Evaluation Protocol=Zero-shot2026.06 | 80.61 | |
| AutoPruneParams=LLaMA-2 13B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 80.58 | |
| BF16Backbone=Llama-2-13B, Quantization=None2026.05 | 80.55 | |
| DenseParams=LLaMA-2 13B, Sparsity=0%, Evaluation protocol=zero-shot2025.11 | 80.52 | |
| CRePEModel=LLaMA-2-7B, Sparsity Pattern=unstructured 50%, Evaluation Protocol=Zero-shot2026.06 | 80.31 | |
| SAGE-PTQBackbone=Qwen-3-8B, Evaluation Protocol=zero-shot2026.06 | 80.28 | |
| GEMQ#Bits=2.5-162026.05 | 80.09 | |
| LoRA-MixerBackbone=LLaMA3-8B, LoRA Rank (r)=322025.06 | 79.37 | |
| MagnitudeBackbone=LLaMA-1, Model Parameters=65B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot, Quantization=8-bit2025.11 | 79.15 | |
| BEAMModel=Qwen1.5-MoE-A2.7B, Beta (β)=0.01, Avg. K=1.562026.05 | 78.32 | |
| DenseBackbone=LLaMA-1, Model Parameters=13B, Sparsity=0%, Evaluation Protocol=zero-shot2025.11 | 77.89 | |
| CRePEModel=LLaMA-2-7B, Sparsity Pattern=2:4, Evaluation Protocol=Zero-shot2026.06 | 77.8 | |
| DenseParams=LLaMA-2 7B, Sparsity=0%, Evaluation protocol=zero-shot2025.11 | 77.74 | |
| BF16Backbone=LLaMA-2-7B, Quantization=None2026.05 | 77.71 | |
| SUBFITModel=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 77.13 | |
| SparseGPTBackbone=LLaMA-1, Model Parameters=13B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 76.97 | |
| AutoPruneBackbone=LLaMA-1, Model Parameters=13B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 76.94 | |
| BEAMModel=DeepSeekV2-Lite, Beta (β)=0.01, Avg. K=2.612026.05 | 76.57 | |
| WandaParams=LLaMA-2 7B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 75.99 | |
| WandaBackbone=LLaMA-1, Model Parameters=13B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 75.9 | |
| Llama-3.2-3B-InstructModel=Llama-3.2-3B-Instruct, Sparsity=Dense2026.06 | 75.63 | |
| AutoPruneParams=LLaMA-2 7B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 75.5 | |
| DeepSeekV2-LiteModel=DeepSeekV2-Lite, Target Top-K (K)=6, Avg. K=6.002026.05 | 75.2 | |
| DenseBackbone=LLaMA-1, Model Parameters=7B, Sparsity=0%, Evaluation Protocol=zero-shot2025.11 | 75.05 | |
| SparseGPTParams=LLaMA-2 7B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 75.02 | |
| LoRABackbone=LLaMA3-8B, LoRA Rank (r)=322025.06 | 74.46 | |
| SAGE-PTQBackbone=LLaMA-3-8B, Evaluation Protocol=zero-shot2026.06 | 74.43 | |
| LLaMA3Size=3.0B, Bit=16.0, Zero-shot=true2025.08 | 73.4 | |
| Qwen1.5-MoEModel=Qwen1.5-MoE-A2.7B, Target Top-K (K)=4, Avg. K=4.002026.05 | 72.63 | |
| AutoPruneBackbone=LLaMA-1, Model Parameters=7B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 72.29 | |
| SparseGPTBackbone=LLaMA-1, Model Parameters=7B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 72.05 | |
| BaseBackbone=LLaMA3-8B2025.06 | 71.25 | |
| WandaBackbone=LLaMA-1, Model Parameters=7B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 71.22 | |
| MagnitudeParams=LLaMA-2 70B, Sparsity=50% unstructured, Evaluation protocol=zero-shot2025.11 | 70.55 | |
| FO (Adam)Base Model=LLaMA-3.2-1B, Optimization Regime=First-Order2025.10 | 69 | |
| FP16Backbone=OPT-6.7B, Evaluation Protocol=zero-shot2026.06 | 68.66 | |
| SUBFITModel=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 68.1 | |
| LGPBackbone=Llama-2-13B, Quantization=W2A16, alpha_1=1e52026.05 | 67.09 | |
| OmniQuantBackbone=Llama-2-13B, Quantization=W2A162026.05 | 66.69 | |
| Entropy-based + SPSequence Reduction=6.3†, FLOPs/byte Reduction=2.72026.05 | 66.3 | |
| ZO Fine-tunerBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 66 | |
| BiLLMBackbone=LLaMA-3-8B, Evaluation Protocol=zero-shot2026.06 | 65.93 | |
| H-Net + SPSequence Reduction=5.3†, FLOPs/byte Reduction=2.42026.05 | 65.9 | |
| Fixed (p = 4) + SPSequence Reduction=4.0, FLOPs/byte Reduction=2.12026.05 | 65.2 | |
| TransformerSize=1.7B, Optimizer=Muon, RoPE on ρ=—2026.05 | 64.92 | |
| SAGE-PTQBackbone=OPT-6.7B, Evaluation Protocol=zero-shot2026.06 | 64.71 | |
| PythiaSize=2.8B, Bit=16.0, Zero-shot=true2025.08 | 64.7 | |
| ParallaxSize=1.7B, Optimizer=Muon, RoPE on ρ=✓2026.05 | 64.59 | |
| Byte-levelSequence Reduction=1.0, FLOPs/byte Reduction=1.02026.05 | 64.5 | |
| H-NetSequence Reduction=4.8, FLOPs/byte Reduction=3.42026.05 | 64.5 | |
| MagnitudeBackbone=LLaMA-1, Model Parameters=30B, Sparsity=50% unstructured, Evaluation Protocol=zero-shot2025.11 | 64.34 | |
| Transformerevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 64.28 | |
| SpaceByte + SPSequence Reduction=6.3, FLOPs/byte Reduction=2.72026.05 | 64.2 | |
| CARVE + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 64.04 | |
| LLaMA3Size=1.3B, Bit=16.0, Zero-shot=true2025.08 | 64 | |
| BiLLMBackbone=Qwen-3-8B, Evaluation Protocol=zero-shot2026.06 | 63.9 | |
| BF16Backbone=Qwen-3-0.6B, Quantization=None2026.05 | 63.82 | |
| Qwen3-0.6B (Pretrained)2026.05 | 63.8 |