Common Sense Reasoning on HellaSwag (acc_n)
95.7Accuracy (acc_n)Always Tell Me The Odds
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn, Training Strategy=+R2025.05 | 95.7 | — | |
| Zamba-7BTraining Tokens (B)=1000, Attention configuration=Hybrid Softmax2025.07 | 76.4 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn2025.05 | 75.5 | — | |
| GPT-4oEvaluation Protocol=0-Shot2025.05 | 75.4 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct2025.05 | 75.3 | — | |
| Llama-3-InstructEvaluation Protocol=Probe, Size=14B2025.05 | 74.9 | — | |
| LLaMA-3-8B-LizardTraining Tokens (B)=0.04, Attention configuration=Linearized (Keep 50% Full Attn.)2025.07 | 73.6 | — | |
| LLaMA-3-8BTraining Tokens (B)=15000, Attention configuration=Softmax2025.07 | 73.1 | — | |
| Mamba2-LLaMA-3Training Tokens (B)=20, Attention configuration=Linearized (Keep 50% Full Attn.)2025.07 | 71.5 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-7B-Instruct2025.05 | 70.2 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-8B-Instruct2025.05 | 67.2 | — | |
| StripedHyena-Nous-7BTraining Tokens (B)=–, Attention configuration=Hybrid Softmax2025.07 | 66.4 | — | |
| Llama-3.2-3BQuantization Strategy=W4A16, Sample size (n)=202026.04 | 65 | — | |
| Llama-3.2-3BQuantization Strategy=W4A8, Sample size (n)=202026.04 | 65 | — | |
| Llama-3.2-3BQuantization Strategy=MCAP Mixed, Sample size (n)=202026.04 | 65 | — | |
| DeepSeek-R1-Distill-Qwen-32BEvaluation Protocol=0-Shot2025.05 | 57.8 | — | |
| Pruner-ZeroBase Model=LLaMA-2 7B, Pruning Strategy=Semi-Structured Pruning, Parametric Budget=2:4, Evaluation Protocol=10-shot2026.06 | 54.7 | — | |
| Llama-3.2-1BQuantization Strategy=W4A16, Sample size (n)=502026.04 | 54 | — | |
| Llama-3.2-1BQuantization Strategy=W4A8, Sample size (n)=502026.04 | 54 | — | |
| Llama-3.2-1BQuantization Strategy=MCAP Mixed, Sample size (n)=502026.04 | 54 | — | |
| DOT-MoEBase Model=LLaMA-2 7B, Pruning Strategy=Semi-Structured Pruning / MoE Conversion, Parametric Budget=2:4, Evaluation Protocol=10-shot2026.06 | 53.9 | — | |
| DISP-LLMBase Model=LLaMA-2 7B, Pruning Strategy=Structured Pruning, Parametric Budget=50%, Evaluation Protocol=10-shot2026.06 | 46.3 | — | |
| ShortGPTBase Model=LLaMA-2 7B, Pruning Strategy=Structured Pruning, Parametric Budget=50%, Evaluation Protocol=10-shot2026.06 | 43.7 | — | |
| SparseGPTBase Model=LLaMA-2 7B, Pruning Strategy=Semi-Structured Pruning, Parametric Budget=2:4, Evaluation Protocol=10-shot2026.06 | 43.3 | — | |
| RoBERTa-LType=Encoder2025.05 | 42 | — | |
| WandaBase Model=LLaMA-2 7B, Pruning Strategy=Semi-Structured Pruning, Parametric Budget=2:4, Evaluation Protocol=10-shot2026.06 | 40.9 | — | |
| Glauber-MN=32026.05 | 40.5 | — | |
| Original (uncompressed)Backbone=Pythia 1.4B, Evaluation Protocol=zero-shot2026.05 | 40.41 | 52.02 | |
| LLM SurgeonBase Model=LLaMA-2 7B, Pruning Strategy=Structured Pruning, Parametric Budget=50%, Evaluation Protocol=10-shot2026.06 | 40.3 | — | |
| GPT-2-M2026.05 | 38.3 | — | |
| Mamba-2Model Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 37.7 | — | |
| Mamba-2 + PoSTModel Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 37.5 | — | |
| Glauber-MN=12026.05 | 37.4 | — | |
| (Gong et al., 2025)-M2026.05 | 37.2 | — | |
| K-OBDBase Model=LLaMA-2 7B, Pruning Strategy=Structured Pruning, Parametric Budget=50%, Evaluation Protocol=10-shot2026.06 | 36.8 | — | |
| SVD-LLM V1Backbone=Pythia 1.4B, Compression Ratio (d/D)=0.25, Evaluation Protocol=zero-shot2026.05 | 34.49 | 43.53 | |
| SliceGPTBase Model=LLaMA-2 7B, Pruning Strategy=Structured Pruning, Parametric Budget=50%, Evaluation Protocol=10-shot2026.06 | 33 | — | |
| SVD-LLM V2Backbone=Pythia 1.4B, Compression Ratio (d/D)=0.25, Evaluation Protocol=zero-shot2026.05 | 32.4 | 40.29 | |
| RWKV-7Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 32.1 | — | |
| RWKV-7 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 32.1 | — | |
| Gated DeltaNetModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 31.9 | — | |
| Gated DeltaNet + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 31.5 | — | |
| SEDD-M2026.05 | 31.5 | — | |
| Mamba-2 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 31.3 | — | |
| Mamba-2Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 31.1 | — | |
| Pure-bundle GBD (λ=0)Backbone=Pythia 1.4B, Compression Ratio (d/D)=0.25, Evaluation Protocol=zero-shot2026.05 | 27.85 | 30.33 | |
| Basis-Sharing-coreBackbone=Pythia 1.4B, Compression Ratio (d/D)=0.25, Evaluation Protocol=zero-shot2026.05 | 26.18 | 26.76 | |
| CARTd=256, Best R=8, zero-shot=true2026.05 | — | 26.49 | |
| CARTd=512, Best R=6, zero-shot=true2026.05 | — | 27.1 | |
| CARTd=768, Best R=6, zero-shot=true2026.05 | — | 27.05 | |
| CARTd=1024, Best R=6, zero-shot=true2026.05 | — | 28.04 | |
| CCQ-Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 44.61 | |
| CCQ-Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 56.82 | |
| CCQ-GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 43.45 | |
| CCQ-GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 59.14 | |
| Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 43.48 | |
| Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 56.66 | |
| GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 41.93 | |
| GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 55.44 | |
| GLA-HedgehogScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 40.48 | |
| GLA-HedgehogScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 55.9 | |
| Mamba2Scale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 44.15 | |
| Mamba2Scale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | — | 57 | |
| TransformerScale=500M, Training tokens=15B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | — | 43.12 | |
| TransformerScale=1.3B, Training tokens=40B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | — | 53.44 |