Language Modeling on 1 Billion Word Benchmark 1.0 (test)
28Test Perplexity (10 epochs)High-Budget MoE Model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| High-Budget MoE Model#Parameters excluding embedding and softmax layers=4371 million, ops/timestep=142.7 million, Training Time (10 epochs)=47 hours, 32 k40s, TFLOPS/GPU=1.562017.01 | 28 | — | |
| Medium-Budget MoE Model#Parameters excluding embedding and softmax layers=4313 million, ops/timestep=33.8 million, Training Time (10 epochs)=17 hours, 32 k40s, TFLOPS/GPU=1.222017.01 | 31.3 | — | |
| Low-Budget MoE Model#Parameters excluding embedding and softmax layers=4303 million, ops/timestep=8.9 million, Training Time (10 epochs)=15 hours, 16 k40s, TFLOPS/GPU=0.742017.01 | 34.1 | — | |
| Best Published Results (Jozefowicz et al., 2016)#Parameters excluding embedding and softmax layers=151 million, ops/timestep=151 million, Training Time (10 epochs)=59 hours, 32 k40s, TFLOPS/GPU=1.092017.01 | 34.7 | 30.6 |