Commonsense Reasoning on Winogrande (HS, Acc)
82.35AccuracyNone
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| NoneModel=Mistral, Mixing ratio (r)=0.12026.03 | 82.35 | 47.75 | |
| NoneModel=Llama3, Mixing ratio (r)=0.12026.03 | 82.22 | 5.25 | |
| BeavertailsModel=Llama3, Mixing ratio (r)=0.12026.03 | 81.98 | 29.18 | |
| GR-SAPModel=Llama3, Mixing ratio (r)=0.12026.03 | 81.64 | 0.68 | |
| AegisModel=Llama3, Mixing ratio (r)=0.12026.03 | 81.61 | 1.42 | |
| AegisModel=Mistral, Mixing ratio (r)=0.12026.03 | 80.77 | 9 | |
| NoneModel=Qwen2.5, Mixing ratio (r)=0.12026.03 | 79.32 | 17.3 | |
| AegisModel=Qwen2.5, Mixing ratio (r)=0.12026.03 | 77.93 | 14.42 | |
| BeavertailsModel=Qwen2.5, Mixing ratio (r)=0.12026.03 | 77.93 | 40.5 | |
| GR-SAPModel=Qwen2.5, Mixing ratio (r)=0.12026.03 | 77.53 | 11.04 | |
| NoneModel=OLMo2, Mixing ratio (r)=0.12026.03 | 76.06 | 0.79 | |
| GR-SAPModel=OLMo2, Mixing ratio (r)=0.12026.03 | 75.19 | 0.45 | |
| BeavertailsModel=OLMo2, Mixing ratio (r)=0.12026.03 | 75.09 | 21.99 | |
| OriginalModel=OLMo2, Mixing ratio (r)=0.12026.03 | 75.06 | 0.5 | |
| AegisModel=OLMo2, Mixing ratio (r)=0.12026.03 | 74.98 | 2.15 | |
| GRINQH-6b-RTNEff. Bits=3.881, Model=Llama-3.1-8B-Instruct2026.06 | 74.66 | — | |
| GRINQH-8b-RTNEff. Bits=3.96, Model=Llama-3.1-8B-Instruct2026.06 | 74.35 | — | |
| MatGPTQ-EP-Mix’n’MatchEff. Bits=4.00, Model=Llama-3.1-8B-Instruct2026.06 | 73.56 | — | |
| BF16-BaselineEff. Bits=16, Model=Llama-3.1-8B-Instruct2026.06 | 73.48 | — | |
| GRINQH-6b-RTNEff. Bits=2.98, Model=Llama-3.1-8B-Instruct2026.06 | 73.16 | — | |
| GR-SAPModel=Mistral, Mixing ratio (r)=0.12026.03 | 73.09 | 12.33 | |
| GRINQH-8b-RTNEff. Bits=2.95, Model=Llama-3.1-8B-Instruct2026.06 | 73.01 | — | |
| MatGPTQ-EP-Mix’n’MatchEff. Bits=3.00, Model=Llama-3.1-8B-Instruct2026.06 | 72.53 | — | |
| GRINQH-4b-GPTQEff. Bits=3.01, Model=Llama-3.1-8B-Instruct2026.06 | 72.53 | — | |
| MatGPTQEff. Bits=4.00, Model=Llama-3.1-8B-Instruct2026.06 | 72.06 | — | |
| MatGPTQEff. Bits=3.00, Model=Llama-3.1-8B-Instruct2026.06 | 71.27 | — | |
| BeavertailsModel=Mistral, Mixing ratio (r)=0.12026.03 | 69.98 | 41.98 | |
| GRINQH-8b-RTNEff. Bits=2.24, Model=Qwen3-8B2026.06 | 67.72 | — | |
| GRINQH-6b-RTNEff. Bits=3.05, Model=Qwen3-8B2026.06 | 67.64 | — | |
| BF16-BaselineEff. Bits=16, Model=Qwen3-8B2026.06 | 67.56 | — | |
| GRINQH-8b-RTNEff. Bits=2.08, Model=Llama-3.1-8B-Instruct2026.06 | 67.4 | — | |
| GRINQH-4b-GPTQEff. Bits=2.09, Model=Qwen3-8B2026.06 | 66.93 | — | |
| GRINQH-4b-GPTQEff. Bits=3.02, Model=Qwen3-8B2026.06 | 66.61 | — | |
| GRINQH-8b-RTNEff. Bits=3.06, Model=Qwen3-8B2026.06 | 65.9 | — | |
| MatGPTQEff. Bits=3.00, Model=Qwen3-8B2026.06 | 64.64 | — | |
| MatGPTQ-EP-Mix’n’MatchEff. Bits=3.00, Model=Qwen3-8B2026.06 | 64.48 | — | |
| Gated KalmaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 64.17 | — | |
| Gated DeltaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 62.35 | — | |
| Mamba2evaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 62.19 | — | |
| Transformerevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 61.72 | — | |
| DeltaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 61.72 | — | |
| CARVE + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 60.71 | — | |
| Muon + PCZero-shot=true, Model=Llama-1B2026.06 | 59.67 | — | |
| GDN-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 58.56 | — | |
| CARVEModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 58.12 | — | |
| Muon baselineZero-shot=true, Model=Llama-1B2026.06 | 58.09 | — | |
| GDN-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.85 | — | |
| KDA + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.77 | — | |
| Mamba-3 SISO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.3 | — | |
| Mamba-3 MIMO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.06 | — | |
| Gated DeltaNet + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.83 | — | |
| Gated DeltaNetModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.75 | — | |
| AdamW + PCZero-shot=true, Model=Llama-1B2026.06 | 56.27 | — | |
| Mamba-3 SISOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.2 | — | |
| Mamba-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.17 | — | |
| TransformerModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.85 | — | |
| Mamba-3 MIMOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.78 | — | |
| KDAModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.72 | — | |
| Mamba-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.33 | — | |
| Gated Linear Attentionevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 54.54 | — | |
| AdamW baselineZero-shot=true, Model=Llama-1B2026.06 | 54.38 | — |