Language Modeling on 100 Billion Word Google News Dataset (test)
38.2Test Perplexity (0.1 epochs)MoE-16384-h
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MoE-16384-hops/timestep (millions)=8.8, #Params excluding embed. & softmax (millions)=17201, Total #Params (billions)=17.3, TFLOPS per GPU (observed)=0.962017.01 | 38.2 | 29.7 | |
| MoE-65536-hops/timestep (millions)=9.2, #Params excluding embed. & softmax (millions)=68791, Total #Params (billions)=68.9, TFLOPS per GPU (observed)=0.722017.01 | 38.2 | 28.9 | |
| MoE-4096-hops/timestep (millions)=8.6, #Params excluding embed. & softmax (millions)=4303.4, Total #Params (billions)=4.4, TFLOPS per GPU (observed)=1.072017.01 | 38.9 | 30.9 | |
| MoE-131072-hops/timestep (millions)=9.7, #Params excluding embed. & softmax (millions)=137577.6, Total #Params (billions)=137.7, TFLOPS per GPU (observed)=0.32017.01 | 39.8 | 29.2 | |
| MoE-1024-hops/timestep (millions)=8.5, #Params excluding embed. & softmax (millions)=1079, Total #Params (billions)=1.2, TFLOPS per GPU (observed)=1.142017.01 | 40.3 | 32.7 | |
| MoE-256-hops/timestep (millions)=8.4, #Params excluding embed. & softmax (millions)=272.9, Total #Params (billions)=0.4, TFLOPS per GPU (observed)=1.112017.01 | 42.8 | 35.3 | |
| MoE-32ops/timestep (millions)=8.4, #Params excluding embed. & softmax (millions)=37.8, Total #Params (billions)=0.1, TFLOPS per GPU (observed)=0.832017.01 | 48.5 | 40.4 | |
| 4xLSTM-512ops/timestep (millions)=8.4, #Params excluding embed. & softmax (millions)=8.4, Total #Params (billions)=0.1, TFLOPS per GPU (observed)=1.232017.01 | 54.5 | 47 | |
| Kneser-Ney 5-gramops/timestep (millions)=0.00001, Total #Params (billions)=762017.01 | 67.1 | 45.3 |