Language Modeling on LAMBADA (PPL, Accuracy)
11.39PPLCCQ-GLA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CCQ-GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 11.39 | 47.82 | |
| GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 12.31 | 47.02 | |
| CCQ-Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 12.52 | 46.13 | |
| Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 13.11 | 45.57 | |
| Mamba2Scale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 13.55 | 45.55 | |
| TransformerScale=1.3B, Training tokens=40B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 14.31 | 45.02 | |
| GLA-HedgehogScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 15.17 | 43.08 | |
| CCQ-Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 22.93 | 38.19 | |
| CCQ-GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 25.77 | 36.74 | |
| Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 29.1 | 33.88 | |
| Mamba2Scale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 29.37 | 34.62 | |
| TransformerScale=500M, Training tokens=15B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 30.47 | 34.04 | |
| GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 32.11 | 34.02 | |
| GLA-HedgehogScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 43.74 | 29.5 |