Zero-shot Word Prediction on LAMBADA
24.16AccuracyYat (pn+α)
Evaluation Results
| Method | Links | |
|---|---|---|
| Yat (pn+α)Parameterization=per-neuron, Scaling=learnable per-channel, Model size=261M, Training tokens=5.2B C4, Evaluation mode=zero-shot2026.05 | 24.16 | |
| Yat (sb+α)Parameterization=shared (b, ε), Scaling=learnable per-channel, Model size=261M, Training tokens=5.2B C4, Evaluation mode=zero-shot2026.05 | 23.8 | |
| Yat (pn+ca)Parameterization=per-neuron, Scaling=constant per-channel, Model size=261M, Training tokens=5.2B C4, Evaluation mode=zero-shot2026.05 | 23.79 | |
| GELUArchitecture=GELU MLP, Model size=261M, Training tokens=5.2B C4, Evaluation mode=zero-shot2026.05 | 23.59 | |
| Yat (sb+ca)Parameterization=shared (b, ε), Scaling=constant per-channel, Model size=261M, Training tokens=5.2B C4, Evaluation mode=zero-shot2026.05 | 23.43 |