Commonsense Reasoning on Winogrande (accuracy)
85.3AccuracyLlama-3-70B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama-3-70B2024.07 | 85.3 | — | |
| Qwen2-72B2024.07 | 85.1 | — | |
| Mixtral-8x22B2024.07 | 85 | — | |
| Yi-1.5-34BArchitecture=Dense, # Act Params=32B, # Params=32B2024.07 | 84.9 | — | |
| Qwen1.5-110B2024.07 | 83.5 | — | |
| Qwen1.5-72B2024.07 | 83 | — | |
| JambaArchitecture=MoE, # Act Params=12B, # Params=52B2024.07 | 82.5 | — | |
| Mixtral-8x7BArchitecture=MoE, # Act Params=12B, # Params=47B2024.07 | 81.9 | — | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 81.5 | — | |
| DeepSeek-R1 (BF16)Model=DeepSeek-R1, Setting=W16A16, Sparsity=0, Evaluation Protocol=Non-generative2025.12 | 79.95 | — | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 79.5 | — | |
| SQ-formatModel=DeepSeek-R1, Setting=W(SQ5)A8, Sparsity=0.875, Evaluation Protocol=Non-generative2025.12 | 79.45 | — | |
| GTCABackbone Model=Llama-3-8B, Evaluation Protocol=5-shot2026.01 | 77.89 | — | |
| Gemma3Parameters=12B2026.03 | 77.74 | — | |
| FP-16Backbone=LLaMA-1-65B, Bit-width=W16A16, Evaluation Protocol=zero-shot2024.02 | 77.11 | — | |
| Llama3.1Parameters=8B2026.03 | 77.11 | — | |
| Direct-JointBackbone Model=Llama-3-8B, Evaluation Protocol=5-shot2026.01 | 76.98 | — | |
| Qwen3Parameters=8B2026.03 | 76.8 | — | |
| BackboneBackbone Model=Llama-3-8B, Evaluation Protocol=5-shot2026.01 | 76.72 | — | |
| Qwen2.5Parameters=7B2026.03 | 76.48 | — | |
| LoopRPTParameters=2.6B2026.03 | 76.47 | — | |
| Vector (row/col)Model=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 76.24 | — | |
| BitDelta (scalar)Model=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 76.09 | — | |
| LoRA-onlyBackbone Model=Llama-3-8B, Evaluation Protocol=5-shot2026.01 | 76.01 | — | |
| OuroParameters=2.6B2026.03 | 75.85 | — | |
| BaselineModel=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 75.61 | — | |
| Traditional MTL2024.05 | 75.5 | — | |
| Individual2024.05 | 75.1 | — | |
| LoRA-onlyBackbone Model=Qwen-2.5-7B, Evaluation Protocol=5-shot2026.01 | 75.1 | — | |
| GTCABackbone Model=Qwen-2.5-7B, Evaluation Protocol=5-shot2026.01 | 74.95 | — | |
| SPINIteration=Iter32025.12 | 74.51 | — | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 74.4 | — | |
| SPINIteration=Iter42025.12 | 74.35 | — | |
| Qwen2-57B-A14BPruning ratio=0%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 74.27 | — | |
| Direct-JointBackbone Model=Qwen-2.5-7B, Evaluation Protocol=5-shot2026.01 | 74.23 | — | |
| Vector (row/col)Model=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 74.19 | — | |
| Zephyr-7BIteration=N/A2025.12 | 74.19 | — | |
| SPINIteration=Iter12025.12 | 74.11 | — | |
| SPACEIteration=Iter32025.12 | 74.11 | — | |
| LLaMA2Parameters=7B, Training Stage=Pretrained comparison2024.01 | 74.03 | — | |
| SPACEIteration=Iter22025.12 | 74.03 | — | |
| Self-MOA aligned model (M_Self-MOA)Model Family=gemma-2-2b-it2026.03 | 74.03 | — | |
| LLAMA PROParameters=8B, Training Stage=Pretrained comparison2024.01 | 73.95 | — | |
| BitDelta (scalar)Model=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 73.95 | — | |
| SPINIteration=Iter22025.12 | 73.95 | — | |
| CD-MoE-SR#L=9, ACT.=2.8B, MEM.=72.5%, Speedup=1.26x, SFT Protocol=w/ lightweight SFT2024.11 | 73.9 | — | |
| SPACEIteration=Iter02025.12 | 73.88 | — | |
| Mistral-7B-v0.3Attack Status=Before2025.11 | 73.87 | — | |
| BaselineModel=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 73.87 | — | |
| BaselineModel=Llama3.1 8B2026.03 | 73.87 | — | |
| SPACEIteration=Iter12025.12 | 73.64 | — | |
| S-SimPOIteration=Iter32025.12 | 73.48 | — | |
| SPINIteration=Iter02025.12 | 73.48 | — | |
| SPACEIteration=Iter42025.12 | 73.48 | — | |
| SpB1.0-7BModel Size=7B2026.04 | 73.48 | — | |
| S-SimPOIteration=Iter22025.12 | 73.4 | — | |
| BitDelta (scalar)Model=Qwen3-14B, Zero-shot protocol=true2025.12 | 73.16 | — | |
| BaselineModel=Qwen3 32B2026.03 | 73.16 | — | |
| S-IPOIteration=Iter32025.12 | 73.09 | — | |
| BackboneBackbone Model=Qwen-2.5-7B, Evaluation Protocol=5-shot2026.01 | 73.01 | — | |
| SEPTQBits=3, Size=30B, Evaluation Protocol=zero-shot2026.04 | 73.01 | — | |
| S-IPOIteration=Iter12025.12 | 72.98 | — | |
| Vector (row/col)Model=Qwen3-14B, Zero-shot protocol=true2025.12 | 72.93 | — | |
| S-SimPOIteration=Iter12025.12 | 72.93 | — | |
| FullBits=16, Size=30B, Evaluation Protocol=zero-shot2026.04 | 72.93 | — | |
| S-SimPOIteration=Iter42025.12 | 72.8 | — | |
| FP-16Backbone=LLaMA-1-30B, Bit-width=W16A16, Evaluation Protocol=zero-shot2024.02 | 72.77 | — | |
| Llama3-8BAttack Status=Before2025.11 | 72.77 | — | |
| BaselineModel=Qwen3-14B, Zero-shot protocol=true2025.12 | 72.77 | — | |
| Gemma3-4BModel Size=4B2026.04 | 72.77 | — | |
| WizardMathParameters=7B, Training Stage=SFT comparison2024.01 | 72.69 | — | |
| S-IPOIteration=Iter42025.12 | 72.69 | — | |
| S-SimPOIteration=Iter02025.12 | 72.69 | — | |
| S-IPOIteration=Iter22025.12 | 72.61 | — | |
| LLAMA PRO INSTRUCTTraining Stage=SFT comparison2024.01 | 72.53 | — | |
| FalconParameters=7B, Training Stage=Pretrained comparison2024.01 | 72.38 | — | |
| Llama3.2-3BModel Size=3B2026.04 | 72.38 | — | |
| Llama2-13BAttack Status=Before2025.11 | 72.3 | — | |
| S-IPOIteration=Iter02025.12 | 72.3 | — | |
| BF16Backbone=Llama-2-13B2026.05 | 72.14 | — | |
| QLoRAModel=Llama-2-7B, Average Bit Budget=3-bit, Zero-shot=true2025.09 | 71.88 | — | |
| GPTQBackbone=LLaMA-1-65B, Bit-width=W2A16, Evaluation Protocol=zero-shot2024.02 | 71.82 | — | |
| DB-LLMBackbone=LLaMA-1-65B, Bit-width=W2A16, Evaluation Protocol=zero-shot2024.02 | 71.82 | — | |
| LLaMA2-ChatParameters=7B, Training Stage=SFT comparison2024.01 | 71.74 | — | |
| H2ONumber of shots=5-shot, KV cache budget=20%, Sparsification Method=Heavy Hitter Oracle2023.06 | 71.67 | — | |
| FullNumber of shots=5-shot, KV cache budget=100%2023.06 | 71.51 | — | |
| Deepseek-V2-LitePruning ratio=0%, Base Model=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 71.51 | — | |
| LLaMAParameters=7B, Training Stage=Pretrained comparison2024.01 | 71.43 | — | |
| 50% KV PredModel=Falcon3 10B2026.03 | 71.42 | — | |
| PB-LLMBackbone=LLaMA-1-65B, Bit-width=W2A16, Evaluation Protocol=zero-shot2024.02 | 71.27 | — | |
| Qwen3Parameters=4B2026.03 | 71.27 | — | |
| BaselineModel=Falcon3 10B2026.03 | 71.19 | — | |
| Qwen2.5Parameters=4B2026.03 | 71.19 | — | |
| MoonlightPruning ratio=0%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 71.11 | — | |
| GPTQBits=3, Size=30B, Evaluation Protocol=zero-shot2026.04 | 71.03 | — | |
| SEPTQBits=2, Size=30B, Evaluation Protocol=zero-shot2026.04 | 70.88 | — | |
| QuIPBits=3, Size=30B, Evaluation Protocol=zero-shot2026.04 | 70.88 | — | |
| RSPruning ratio=50%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 70.8 | — | |
| Qwen3-30B-A3BPruning ratio=0%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 70.64 | — | |
| Qwen3-4B-BaseModel Size=4B, Training Strategy=Base2026.04 | 70.48 | — |