Sentence Completion on HellaSwag
87.5AccuracyFalcon-180B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Falcon-180Bprotocol=1-shot2023.11 | 87.5 | — | — | — | |
| PaLM-2 Lprotocol=1-shot2023.11 | 86.8 | — | — | — | |
| Full-precisionModel=Mixtral 8x22B, Avg. bits/exp.=16 (FP), Memory (GB)=281.2, Evaluation protocol=zero-shot2026.04 | 84.5 | — | — | — | |
| HICDBackbone=LLaMA2-7b2025.03 | 84.33 | — | — | — | |
| HICDBackbone=LLaMA-7b2025.03 | 84.23 | — | — | — | |
| PaLM-2 Mprotocol=1-shot2023.11 | 84 | — | — | — | |
| Full-precisionModel=Mixtral 8x7B, Avg. bits/exp.=16 (FP), Memory (GB)=96.8, Evaluation protocol=zero-shot2026.04 | 83.99 | — | — | — | |
| PaLMprotocol=1-shot2023.11 | 83.6 | — | — | — | |
| COVERCALModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 82.4 | — | — | — | |
| PaLM-2 Sprotocol=1-shot2023.11 | 82 | — | — | — | |
| Max-ActVarModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 81.7 | — | — | — | |
| UniformModel=Mixtral 8x7B, Avg. bits/exp.=3, Memory (GB)=19.3, Evaluation protocol=zero-shot2026.04 | 81.51 | — | — | — | |
| RandomModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 81.3 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.75, Memory (GB)=17.7, Evaluation protocol=zero-shot2026.04 | 81.15 | — | — | — | |
| MistralSize=7B, Tokens=?, Shot(s)=0, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 81.1 | — | — | — | |
| COVERCALModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 81.1 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.75, Memory (GB)=17.7, Evaluation protocol=zero-shot2026.04 | 81.05 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.625, Memory (GB)=16.9, Evaluation protocol=zero-shot2026.04 | 80.57 | — | — | — | |
| GemmaSize=7B, Tokens=6T, Shot(s)=0, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 80.5 | — | — | — | |
| Max-ActVarModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 80.3 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.625, Memory (GB)=16.9, Evaluation protocol=zero-shot2026.04 | 80.04 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=16.1, Evaluation protocol=zero-shot2026.04 | 80.02 | — | — | — | |
| RandomModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 79.9 | — | — | — | |
| BaselineModel Architecture=Qwen3-30B-A3B2026.05 | 79.6 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=16.1, Evaluation protocol=zero-shot2026.04 | 79.36 | — | — | — | |
| NHA-Qwen3-30BA3BModel=NHA-Qwen3-30BA3B, Architecture Type=Native Hybrid Attention, Full-attention layers=32025.10 | 79.1 | — | — | — | |
| GPT-3 (175B)Parameters=175B2023.02 | 78.9 | — | — | — | |
| GPT-3Model Size=175B, Evaluation Protocol=Zero-shot2022.10 | 78.9 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.375, Memory (GB)=15.3, Evaluation protocol=zero-shot2026.04 | 78.69 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.375, Memory (GB)=15.3, Evaluation protocol=zero-shot2026.04 | 78.54 | — | — | — | |
| AlpacaBackbone=LLaMA-7b2025.03 | 78.49 | — | — | — | |
| VanillaBackbone=LLaMA2-7b2025.03 | 78.32 | — | — | — | |
| BaselineModel=OLMoE, Weight Bits=16, Zero-shot=true2026.04 | 78.21 | — | — | — | |
| BaselineBackbone=Qwen1.5-MoE, Weight Bitwidth=162026.04 | 78.21 | — | — | — | |
| TORQModel Architecture=Qwen3-30B-A3B2026.05 | 78.12 | — | — | — | |
| HessianModel=Mixtral 8x7B, Avg. bits/exp.=2.5, Memory (GB)=17.0, Evaluation protocol=zero-shot2026.04 | 78.05 | — | — | — | |
| BaselineType=None, Storage=31.41GB, Model Family=DeepSeek-V2-Lite, Evaluation Protocol=Zero-shot2026.06 | 77.98 | — | — | — | |
| GSASize=7B, Tokens=+100B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 77.9 | — | — | — | |
| MambaSize=7B, Tokens=1.2T, Shot(s)=0, Training Strategy=trained from scratch2024.09 | 77.8 | — | — | — | |
| BaselineModel=Qwen3-MoE, Weight Bits=16, Zero-shot=true2026.04 | 77.64 | — | — | — | |
| Qwen3-30BA3BModel=Qwen3-30BA3B, Architecture Type=Transformer2025.10 | 77.63 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.25, Memory (GB)=14.5, Evaluation protocol=zero-shot2026.04 | 77.62 | — | — | — | |
| VanillaBackbone=LLaMA-7b2025.03 | 77.61 | — | — | — | |
| COVERCALModel=LLaMA-2-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 77.6 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.25, Memory (GB)=14.5, Evaluation protocol=zero-shot2026.04 | 77.46 | — | — | — | |
| MxMoEBackbone=Mixtral, Weight Bitwidth=2.252026.04 | 77.44 | — | — | — | |
| BaselineType=None, Storage=32.7GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 77.43 | — | — | — | |
| UNPRUNEDACT.=2.8B, MEM.=100%, Speedup=1.0x, SFT Protocol=w/o SFT2024.11 | 77.4 | — | — | — | |
| BaselineBackbone=Mixtral, Weight Bitwidth=162026.04 | 77.23 | — | — | — | |
| OSTQuantModel Architecture=Qwen3-30B-A3B2026.05 | 77.23 | — | — | — | |
| WandaType=P25% Q4b, Storage=7.70GB, Model Family=DeepSeek-V2-Lite, Evaluation Protocol=Zero-shot2026.06 | 77.12 | — | — | — | |
| SUPRASize=7B, Tokens=+100B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 77.1 | — | — | — | |
| Max-ActVarModel=LLaMA-2-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 76.9 | — | — | — | |
| TWLABackbone=Qwen3-32B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 76.89 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.125, Memory (GB)=13.8, Evaluation protocol=zero-shot2026.04 | 76.76 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.125, Memory (GB)=13.8, Evaluation protocol=zero-shot2026.04 | 76.56 | — | — | — | |
| GSASize=7B, Tokens=+20B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 76.5 | — | — | — | |
| RandomModel=LLaMA-2-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 76.5 | — | — | — | |
| HessianModel=Mixtral 8x7B, Avg. bits/exp.=2.25, Memory (GB)=15.3, Evaluation protocol=zero-shot2026.04 | 76.38 | — | — | — | |
| OursQType=P25% Q4b, Storage=6.00GB, Model Family=DeepSeek-V2-Lite, Evaluation Protocol=Zero-shot2026.06 | 76.25 | — | — | — | |
| LLaMA-7Bzero-shot=true, parameters=7B2023.12 | 76.1 | — | — | — | |
| Llama2Size=7B, Tokens=2T, Shot(s)=0, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 76 | — | — | — | |
| GLASize=7B, Tokens=+20B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 75.9 | — | — | — | |
| EAC-MoEType=P11% Q3.03b, Storage=7.19GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 75.55 | — | — | — | |
| RWKV6Size=7B, Tokens=1.4T, Shot(s)=0, Training Strategy=trained from scratch2024.09 | 75.2 | — | — | — | |
| DoLaBackbone=LLaMA-7b2025.03 | 75.17 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=2.0, Memory (GB)=13.1, Evaluation protocol=zero-shot2026.04 | 74.88 | — | — | — | |
| SUPRASize=7B, Tokens=+20B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 74.8 | — | — | — | |
| FPBackbone=LLADA-1.5-8B, Weight Bits=Full Precision, Activation Bits=Full Precision2026.06 | 74.7 | — | — | — | |
| Attribution-guided and Coverage-Maximized Expert-wise PruningType=P50%, Storage=16.57GB, Model Family=DeepSeek-V2-Lite, Evaluation Protocol=Zero-shot2026.06 | 74.59 | — | — | — | |
| He et al.Type=P25% Q4b, Storage=7.70GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 74.5 | — | — | — | |
| UniformModel=Mixtral 8x22B, Avg. bits/exp.=3, Memory (GB)=57.5, Evaluation protocol=zero-shot2026.04 | 74.23 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=2.0, Memory (GB)=13.1, Evaluation protocol=zero-shot2026.04 | 74.17 | — | — | — | |
| OursQType=P25% Q4b, Storage=6.26GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 73.77 | — | — | — | |
| BaselineModel=DeepSeekV2-Lite, Weight Bits=16, Zero-shot=true2026.04 | 73.55 | — | — | — | |
| FPBackbone=DREAM-7B, Weight Bits=Full Precision, Activation Bits=Full Precision2026.06 | 73.3 | — | — | — | |
| UniformModel=Mixtral 8x7B, Avg. bits/exp.=2, Memory (GB)=13.1, Evaluation protocol=zero-shot2026.04 | 72.93 | — | — | — | |
| RetNetSize=7B, Tokens=+20B, Shot(s)=0, Training Strategy=finetuned from Mistral 7B2024.09 | 72.9 | — | — | — | |
| EAC-MoEType=P38% Q3.03b, Storage=4.47GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 72.86 | — | — | — | |
| TWLABackbone=LLaMA2-13B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 72.59 | — | — | — | |
| STaR-QuantBackbone=LLADA-1.5-8B, Weight Bits=4, Activation Bits=42026.06 | 72.43 | — | — | — | |
| HessianModel=Mixtral 8x7B, Avg. bits/exp.=2.0, Memory (GB)=13.6, Evaluation protocol=zero-shot2026.04 | 71.9 | — | — | — | |
| TWLABackbone=LLaMA2-70B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 71.87 | — | — | — | |
| STaR-QuantBackbone=DREAM-7B, Weight Bits=4, Activation Bits=42026.06 | 71.32 | — | — | — | |
| QuaRotModel Architecture=Qwen3-30B-A3B2026.05 | 71.14 | — | — | — | |
| TWLABackbone=Qwen3-14B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 70.98 | — | — | — | |
| PMQModel=Mixtral 8x7B, Avg. bits/exp.=1.75, Memory (GB)=11.7, Evaluation protocol=zero-shot2026.04 | 70.85 | — | — | — | |
| Attribution-guided and Coverage-Maximized Expert-wise PruningType=P50%, Storage=17.34GB, Model Family=Deepseek-MoE-16B, Evaluation Protocol=Zero-shot2026.06 | 70.5 | — | — | — | |
| DLLMQuant++Backbone=DREAM-7B, Weight Bits=4, Activation Bits=42026.06 | 70.14 | — | — | — | |
| DLLMQuant+Backbone=LLADA-1.5-8B, Weight Bits=4, Activation Bits=42026.06 | 70.09 | — | — | — | |
| NoWagBackbone=Mixtral, Weight Bitwidth=2.082026.04 | 70.08 | — | — | — | |
| Router norm + Max varModel=Mixtral 8x7B, Avg. bits/exp.=1.75, Memory (GB)=11.7, Evaluation protocol=zero-shot2026.04 | 70.03 | — | — | — | |
| DoLaBackbone=LLaMA2-7b2025.03 | 69.25 | — | — | — | |
| DLLMQuant++Backbone=LLADA-1.5-8B, Weight Bits=4, Activation Bits=42026.06 | 69.2 | — | — | — | |
| PMQModel=Mixtral 8x22B, Avg. bits/exp.=2.5, Memory (GB)=46.7, Evaluation protocol=zero-shot2026.04 | 68.64 | — | — | — | |
| GlowQ-SZero-shot=true, Model=Vicuna-13B2026.03 | 67.67 | — | — | — | |
| FP16Rank=-, Zero-shot=true, Model=Vicuna-13B2026.03 | 67.33 | — | — | — | |
| QERAZero-shot=true, Model=Vicuna-13B2026.03 | 67.33 | — | — | — | |
| GlowQZero-shot=true, Model=Vicuna-13B2026.03 | 67.33 | — | — | — | |
| QuaRotBackbone=DREAM-7B, Weight Bits=4, Activation Bits=42026.06 | 67.13 | — | — | — |