Long-range Next-token prediction on PG-19 long-context
101.09Perplexity (PPL)Yat (sb+α)
Evaluation Results
| Method | Links | |
|---|---|---|
| Yat (sb+α)Parameterization=shared (b, ε), Scaling=learnable per-channel, Evaluation strategy=sliding-window, Context length=1024 tokens2026.05 | 101.09 | |
| Yat (pn+α)Parameterization=per-neuron, Scaling=learnable per-channel, Evaluation strategy=sliding-window, Context length=1024 tokens2026.05 | 104.2 | |
| Yat (sb+ca)Parameterization=shared (b, ε), Scaling=constant per-channel, Evaluation strategy=sliding-window, Context length=1024 tokens2026.05 | 106.51 | |
| Yat (pn+ca)Parameterization=per-neuron, Scaling=constant per-channel, Evaluation strategy=sliding-window, Context length=1024 tokens2026.05 | 108.62 | |
| GELUArchitecture=GELU MLP, Evaluation strategy=sliding-window, Context length=1024 tokens2026.05 | 114.08 |