Language Modeling on One Billion Word Benchmark (val)
26.6PerplexityDynamicConv
Evaluation Results
| Method | Links | |
|---|---|---|
| DynamicConvParameters=339M2019.01 | 26.6 | |
| Self-attention baselineParameters=331M2019.01 | 26.67 | |
| MoSParameters=113M2017.11 | 38.01 | |
| SoftmaxParameters=119M2017.11 | 43.86 | |
| Glauber-UL2-LZero-shot=true, Model size=Large, N value=3, Note=upper bound2026.05 | 44.12 | |
| GPT-2-LZero-shot=true, Model size=Large2026.05 | 44.58 | |
| Glauber-UL2-LZero-shot=true, Model size=Large, N value=1, Note=upper bound2026.05 | 47.62 | |
| Glauber-UL2-MZero-shot=true, Model size=Medium, N value=3, Note=upper bound2026.05 | 52.18 | |
| GPT-2-MZero-shot=true, Model size=Medium2026.05 | 55.72 | |
| Glauber-UL2-MZero-shot=true, Model size=Medium, N value=1, Note=upper bound2026.05 | 56.12 | |
| SEDD-MZero-shot=true, Model size=Medium, Note=upper bound2026.05 | 61.19 |