Language Modeling on One Billion Word Benchmark
35.1PerplexityBIGLSTM baseline
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| BIGLSTM baselineNum of RNN parameters=151060480, Training duration=1 week, Hardware=1x DGX Station (4x Tesla V100)2017.03 | 35.1 | — | — | |
| BIG G-LSTM G-2Num of RNN parameters=83951616, Groups=2, Training duration=1 week, Hardware=1x DGX Station (4x Tesla V100)2017.03 | 36 | — | — | |
| BIG F-LSTM F512Num of RNN parameters=52494336, Intermediate rank=512, Training duration=1 week, Hardware=1x DGX Station (4x Tesla V100)2017.03 | 36.3 | — | — | |
| BIG G-LSTM G-8Num of RNN parameters=33619968, Groups=8, Training duration=1 week, Hardware=1x DGX Station (4x Tesla V100)2017.03 | 39.4 | — | — | |
| BIG G-LSTM G-4Num of RNN parameters=50397184, Groups=4, Training duration=1 week, Hardware=1x DGX Station (4x Tesla V100)2017.03 | 40.6 | — | — | |
| GPT-2 1.5Bzero-shot=true, parameters=1.5B2023.05 | 54.09 | — | — | |
| GPT-2 762Mzero-shot=true, parameters=762M2023.05 | 59.48 | — | — | |
| GPT-2 345Mzero-shot=true, parameters=345M2023.05 | 67.34 | — | — | |
| Plaid 1Bzero-shot=true, parameters=1.3B2023.05 | 77.64 | — | — | |
| GPT-2 124Mzero-shot=true, parameters=124M2023.05 | 87.85 | — | — |