Question Answering on ARC Easy (ACC)
90.48AccuracyIT-Prun
Evaluation Results
| Method | Links | |
|---|---|---|
| IT-PrunModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 90.48 | |
| IT-PrunModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.85 | |
| IT-PrunModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 89.68 | |
| LeanQuantModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 88.3 | |
| OSAQ+GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 88.3 | |
| GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 88.2 | |
| ShortGPTModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 88.09 | |
| ShortGPTModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 87.54 | |
| ShortGPTModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 86.53 | |
| LeanQuantModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 85.1 | |
| OSAQ+GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 85 | |
| GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 84.6 | |
| Teacher (DeepSeek-V2-Lite)2026.05 | 84.4 | |
| BBoxERBase Model=Llama3.1-8B, Budget (b)=300, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 83.16 | |
| BBoxERBase Model=Llama3.1-8B, Budget (b)=150, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 83.05 | |
| Llama3.1-8B2025.07 | 83.04 | |
| MKAModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 82.95 | |
| Baseline (Dense)Sparsity=0.0%, Zero-shot=true2026.05 | 82.15 | |
| SLEBModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 81.23 | |
| OriginalBackbone=LLaMA3-8B, Pruning Ratio=0%2026.05 | 81.02 | |
| BBoxERBase Model=Llama3.1-8B-Instruct, Budget (b)=300, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 79.64 | |
| Llama3.1-8B-Instruct2025.07 | 79.62 | |
| MKAModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 79.34 | |
| BaselineModel Architecture=Qwen3-30B-A3B2026.05 | 79.25 | |
| BBoxERBase Model=Llama3.1-8B-Instruct, Budget (b)=150, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 79.16 | |
| MKAModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 79.12 | |
| HARPBackbone=Llama 2 70B, Target Bits=3, BPP=3.04, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 78.6 | |
| QuIP# (RHT)Backbone=Llama 2 70B, Target Bits=3, BPP=3.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 78.4 | |
| HARPBackbone=Llama 2 70B, Target Bits=4, BPP=4.04, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 78.2 | |
| QuIP# (RHT)Backbone=Llama 2 70B, Target Bits=4, BPP=4.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 78.1 | |
| HARPBackbone=Llama 2 70B, Target Bits=2, BPP=2.04, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 77.8 | |
| FP16Backbone=Llama 2 70B, Target Bits=16, BPP=16.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 77.7 | |
| BF16Backbone=Llama-2-13B2026.05 | 77.44 | |
| TALEModel=LLaMA-2-13B, Sparsity=10%, Evaluation Protocol=0-shot2025.10 | 77.3 | |
| TORQModel Architecture=Qwen3-30B-A3B2026.05 | 77.1 | |
| QuIP# (RHT)Backbone=Llama 2 70B, Target Bits=2, BPP=2.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 76.9 | |
| LLAMA-2-7BBase Model=LLaMA-2-7B, Sparsity Pattern=Dense, Evaluation Protocol=Zero-shot2025.06 | 76.35 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 76.2 | |
| OSTQuantModel Architecture=Qwen3-30B-A3B2026.05 | 75.84 | |
| Oryx-TM (Transformer)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 75.7 | |
| Oryx-TG (Transformer)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 75.4 | |
| TALEModel=LLaMA-2-13B, Sparsity=25%, Evaluation Protocol=0-shot2025.10 | 75.3 | |
| Oryx-TM (Mamba-2)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 75.3 | |
| DEEPSEEK-7BBase Model=DeepSeek-7B, Sparsity Pattern=Dense, Evaluation Protocol=Zero-shot2025.06 | 75.25 | |
| Gated DeltaNetParameter Scale=1.4B, Family=Baseline2026.05 | 75 | |
| Accuracy (Arc-E)Backbone=LLaMA3-8B, Pruning Ratio=25%, Calibration Dataset=ARC-E2026.05 | 74.96 | |
| HARPBackbone=Llama 2 13B, Target Bits=4, BPP=4.05, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 74.8 | |
| BF16Backbone=LLaMA-2-7B2026.05 | 74.54 | |
| Mamba-2Parameter Scale=1.4B, Family=Baseline2026.05 | 74.5 | |
| QuIP# (RHT)Backbone=Llama 2 13B, Target Bits=4, BPP=4.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 74.5 | |
| BLTNumber of Parameters=3B2026.05 | 74.33 | |
| Oryx-TM (Transformer)Parameter Scale=810M, Family=Oryx-TM2026.05 | 73.7 | |
| Oryx-TM (Mamba-2)Parameter Scale=810M, Family=Oryx-TM2026.05 | 73.7 | |
| TransformerParameter Scale=1.4B, Family=Baseline2026.05 | 73.6 | |
| Gated DeltaNetParameter Scale=810M, Family=Baseline2026.05 | 73.3 | |
| FP16Backbone=Llama 2 13B, Target Bits=16, BPP=16.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 73.3 | |
| ParallaxSize=1.7B, Optimizer=Muon, RoPE on ρ=✓2026.05 | 73.27 | |
| BaselineModel=LLaMA-2-13B, Sparsity=0%, Evaluation Protocol=0-shot2025.10 | 73 | |
| Mamba-2Parameter Scale=810M, Family=Baseline2026.05 | 72.8 | |
| BBoxERBase Model=Qwen-2.5-3B-instruct, Budget (b)=300, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 72.62 | |
| BBoxERBase Model=Qwen-2.5-3B-instruct, Budget (b)=150, Optimization Algorithm=D-CMA, Update Type=normalization layers2025.07 | 72.57 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=810M, Family=Oryx-TG2026.05 | 72.4 | |
| SLEBModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 72.39 | |
| BLT-DNumber of Parameters=3B, Block size=42026.05 | 72.39 | |
| Qwen-2.5-3B-instruct2025.07 | 72.39 | |
| HARPBackbone=Llama 2 13B, Target Bits=3, BPP=3.05, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 71.9 | |
| Accuracy (Ours)Backbone=LLaMA3-8B, Pruning Ratio=25%, Calibration Dataset=Diverse mix2026.05 | 71.68 | |
| QuIP# (RHT)Backbone=Llama 2 13B, Target Bits=3, BPP=3.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 71.6 | |
| Oryx-TG (Transformer)Parameter Scale=810M, Family=Oryx-TG2026.05 | 71.5 | |
| TransformerParameter Scale=810M, Family=Baseline2026.05 | 71.3 | |
| Oryx-TM (Transformer)Parameter Scale=380M, Family=Oryx-TM2026.05 | 71 | |
| BLT-DNumber of Parameters=3B, Block size=82026.05 | 70.95 | |
| QuIP# (RHT)Backbone=Llama 2 7B, Target Bits=4, BPP=4.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 70.6 | |
| Oryx-TM (Mamba-2)Parameter Scale=380M, Family=Oryx-TM2026.05 | 70.5 | |
| QuaRotModel Architecture=Qwen3-30B-A3B2026.05 | 70.45 | |
| ParallaxSize=1.7B, Optimizer=Muon, RoPE on ρ=✗2026.05 | 70.33 | |
| Slice-GPTBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 70.28 | |
| SPARSEGPTBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 70.2 | |
| TransformerSize=1.7B, Optimizer=Muon, RoPE on ρ=—2026.05 | 69.53 | |
| MaskProBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 69.51 | |
| FP16Backbone=Llama 2 7B, Target Bits=16, BPP=16.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 69.3 | |
| QuIP# (RHT)Backbone=Llama 2 7B, Target Bits=3, BPP=3.00, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 69.3 | |
| PRUNER-ZEROBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 69.23 | |
| HARPBackbone=Llama 2 7B, Target Bits=4, BPP=4.11, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 69.2 | |
| HARPBackbone=Llama 2 7B, Target Bits=3, BPP=3.11, Evaluation Protocol=Zero-shot, Evaluation Harness=lm_eval2026.05 | 69.1 | |
| MaskProBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 68.64 | |
| BF16Model Size=1.3B, Precision=BF16, Zero-shot=true2025.01 | 68.6 | |
| Gated DeltaNetParameter Scale=380M, Family=Baseline2026.05 | 68.6 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=380M, Family=Oryx-TG2026.05 | 68.6 | |
| Mamba-2Parameter Scale=380M, Family=Baseline2026.05 | 68.5 | |
| WANDABase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 68.48 | |
| TransformerParameter Scale=380M, Family=Baseline2026.05 | 68.2 | |
| Wanda-StructSparsity=30.0%, Zero-shot=true2026.05 | 68.18 | |
| PRUNER-ZEROBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 68.18 | |
| GBLMBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 68.18 | |
| SPARSEGPTBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 68.15 | |
| TaylorBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 67.97 | |
| FP4Model Size=7B, Precision=FP4, Zero-shot=true2025.01 | 67.97 | |
| BF16Model Size=13B, Precision=BF16, Zero-shot=true2025.01 | 67.97 | |
| FP4Model Size=13B, Precision=FP4, Zero-shot=true2025.01 | 67.89 |