Common Sense Reasoning on HellaSwag (0-shot)
84.4Accuracyfp16
Evaluation Results
| Method | Links | |
|---|---|---|
| fp16Size=70B2026.02 | 84.4 | |
| OmniQuantSize=70B2026.02 | 83.1 | |
| AstroSize=70B2026.02 | 81.7 | |
| MagRSize=70B2026.02 | 81.1 | |
| fp16Size=13B2026.02 | 79.7 | |
| FP16q (mean bit precision)=16.000, tok/s (throughput)=103.8, Evaluation Protocol=0-shot2025.12 | 79.2 | |
| OmniQuantSize=13B2026.02 | 76.9 | |
| fp16Size=7B2026.02 | 76.2 | |
| AstroSize=13B2026.02 | 76.2 | |
| AQLM-1x16 + PV-Tuningq (mean bit precision)=2.213, tok/s (throughput)=49.0, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 74.33 | |
| MagRSize=13B2026.02 | 74.2 | |
| CodeGEMM-m1v4g128 + PV-Tuningq (mean bit precision)=2.126, tok/s (throughput)=228.3, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 73.85 | |
| OmniQuantSize=7B2026.02 | 73.4 | |
| CodeGEMM-m2v8g128 + PV-Tuningq (mean bit precision)=2.127, tok/s (throughput)=214.4, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 73.02 | |
| FP16Bits=–, Backbone=LLAMA-2-7B2026.04 | 72.96 | |
| AQLM-2x8 + PV-Tuningq (mean bit precision)=2.005, tok/s (throughput)=124.5, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 72.43 | |
| AstroSize=7B2026.02 | 72.2 | |
| AQLM-1x16q (mean bit precision)=2.213, tok/s (throughput)=49.0, Evaluation Protocol=0-shot2025.12 | 70.21 | |
| MagRSize=7B2026.02 | 70.2 | |
| TernaryLLM-8B+KD#Tokens=1T, Quantization=W1.58A162026.06 | 63.9 | |
| CodeGEMM-m1v4g128q (mean bit precision)=2.126, tok/s (throughput)=228.3, Evaluation Protocol=0-shot2025.12 | 63.07 | |
| CodeGEMM-m2v8g128q (mean bit precision)=2.127, tok/s (throughput)=214.4, Evaluation Protocol=0-shot2025.12 | 62.84 | |
| BitNetV1-7B#Tokens=100B, Quantization=W1.58A162026.06 | 61.49 | |
| AQLM-2x8q (mean bit precision)=2.005, tok/s (throughput)=124.5, Evaluation Protocol=0-shot2025.12 | 61.4 | |
| Qwen3-8B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 59.57 | |
| LBLLMBits=W(1+1)A4, Backbone=LLAMA-2-7B2026.04 | 58.54 | |
| Nirvana (Ours)Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 58.25 | |
| FBI-LLMBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 57.7 | |
| Nirvana-noTriggerSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 57.43 | |
| Gated DeltaNet-H2Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 56.88 | |
| Gated DeltaNet-H1Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 56.53 | |
| Gated DeltaNetSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 55.76 | |
| Mamba2Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 55.67 | |
| Qwen3-4B + CAT-Q#Tokens=2M, Quantization=W1.58A162026.06 | 55.01 | |
| SambaSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 53.42 | |
| Qwen3-4B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 53.24 | |
| MambaSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 52.91 | |
| OnebitBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 51.54 | |
| DeltaNetSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 50.93 | |
| Transformer++Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 50.23 | |
| HGRN2Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 49.53 | |
| RetNetSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 49.16 | |
| TriLM-3.9B#Tokens=300B, Quantization=W1.58A162026.06 | 48.3 | |
| Qwen3-1.7B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 47.77 | |
| TequilaLLM-3B#Tokens=10B, Quantization=W1.58A162026.06 | 46.4 | |
| BitNetV1-3.9B#Tokens=100B, Quantization=W1.58A162026.06 | 44.2 | |
| FlexRound-q2g128q (mean bit precision)=2.125, tok/s (throughput)=205.3, Evaluation Protocol=0-shot2025.12 | 43.78 | |
| TriLM-1.5B#Tokens=300B, Quantization=W1.58A162026.06 | 40.9 | |
| BitNetV1-1.3B#Tokens=100B, Quantization=W1.58A162026.06 | 37.7 |