Zero-shot Commonsense Reasoning on PIQA zero-shot
76.93AccuracyFP16
Evaluation Results
| Method | Links | |
|---|---|---|
| FP16Bits=–, Backbone=LLAMA-2-7B2026.04 | 76.93 | |
| TriLM-3.9B#Tokens=300B, Quantization=W1.58A162026.06 | 74.44 | |
| BitNetV1-7B#Tokens=100B, Quantization=W1.58A162026.06 | 74.37 | |
| TequilaLLM-3B#Tokens=10B, Quantization=W1.58A162026.06 | 73.9 | |
| TernaryLLM-8B+KD#Tokens=1T, Quantization=W1.58A162026.06 | 73.7 | |
| Qwen3-8B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 73.45 | |
| BitNetV1-3.9B#Tokens=100B, Quantization=W1.58A162026.06 | 73.2 | |
| FBI-LLMBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 72.6 | |
| Qwen3-4B + CAT-Q#Tokens=2M, Quantization=W1.58A162026.06 | 71.55 | |
| LBLLMBits=W(1+1)A4, Backbone=LLAMA-2-7B2026.04 | 70.97 | |
| Original EndpointModel=1B, Zero-shot evaluation=true2026.06 | 70.95 | |
| Qwen3-4B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 70.62 | |
| TriLM-1.5B#Tokens=300B, Quantization=W1.58A162026.06 | 70.3 | |
| BitNetV1-1.3B#Tokens=100B, Quantization=W1.58A162026.06 | 68.8 | |
| Qwen3-1.7B + CAT-Q#Tokens=1M, Quantization=W1.58A162026.06 | 68.4 | |
| OnebitBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 68.01 | |
| Original PythiaModel=410M, Zero-shot evaluation=true2026.06 | 66.43 | |
| Yat (pn+α)Parameterization=per-neuron, Scaling=learnable per-channel, Evaluation mode=zero-shot2026.05 | 62.44 | |
| Yat (sb+α)Parameterization=shared (b, ε), Scaling=learnable per-channel, Evaluation mode=zero-shot2026.05 | 62.37 | |
| GELUArchitecture=GELU MLP, Evaluation mode=zero-shot2026.05 | 61.77 | |
| Yat (sb+ca)Parameterization=shared (b, ε), Scaling=constant per-channel, Evaluation mode=zero-shot2026.05 | 61.5 | |
| Yat (pn+ca)Parameterization=per-neuron, Scaling=constant per-channel, Evaluation mode=zero-shot2026.05 | 61.26 | |
| Dual Learned MatchingModel=1B, Zero-shot evaluation=true2026.06 | 56.42 | |
| Learned MatchingModel=1B, Zero-shot evaluation=true2026.06 | 55.98 | |
| Dual Learned MatchingModel=410M, Zero-shot evaluation=true2026.06 | 55.71 | |
| Learned MatchingModel=410M, Zero-shot evaluation=true2026.06 | 55.44 | |
| Raw InterpolationModel=410M, Zero-shot evaluation=true2026.06 | 55.39 | |
| Raw InterpolationModel=1B, Zero-shot evaluation=true2026.06 | 54.35 |