Masked Language Modeling on WikiText-103 (train)
6.6188Training LossGEM (N = 1)
Evaluation Results
| Method | Links | |
|---|---|---|
| GEM (N = 1)Architecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=19,427 tok/s, Latency=42.51 ms/step2026.04 | 6.6188 | |
| GEM (N = 2)Architecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=20,022 tok/s, Latency=41.16 ms/step2026.04 | 6.6267 | |
| GELU (tanh)Architecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=19,998 tok/s, Latency=41.16 ms/step2026.04 | 6.628 | |
| GELUArchitecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=19,142 tok/s, Latency=43.03 ms/step2026.04 | 6.6341 | |
| ReLUArchitecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=18,763 tok/s, Latency=43.71 ms/step2026.04 | 6.6413 | |
| SiLU/SwishArchitecture=BERT-small, Number of parameters=28.6M, Training steps=3,000, Seeds=3, Throughput=19,359 tok/s, Latency=42.55 ms/step2026.04 | 6.6436 |