Common-sense Reasoning on PIQA, HellaSwag, WinoGrande, ARC-e, ARC-c, SIQA, and BoolQ
56.24Average AccuracyCCQ-Gated DeltaNet
Evaluation Results
| Method | Links | |
|---|---|---|
| CCQ-Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 56.24 | |
| CCQ-GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 56.1 | |
| GLA-HedgehogScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.93 | |
| GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.75 | |
| Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.64 | |
| Mamba2Scale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.17 | |
| TransformerScale=1.3B, Training tokens=40B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 54.6 | |
| CCQ-Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 49.87 | |
| CCQ-GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 49.73 | |
| TransformerScale=500M, Training tokens=15B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 49.44 | |
| Mamba2Scale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 49.17 | |
| Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 48.98 | |
| GLA-HedgehogScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 48.77 | |
| GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 48.54 |