Language Modeling on OpenWebText (Training Loss)
0.0028Final LossPACI
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PACIBatch=256, Accumulation Factor (a)=162026.06 | 0.0028 | 0.0018 | |
| 1F1B-FlushBatch=2562026.06 | 0.0028 | 0.0053 | |
| PACIBatch=256, Accumulation Factor (a)=82026.06 | 0.0028 | 0.003 | |
| 1F1B-FlushBatch=1282026.06 | 0.0028 | 0.0011 | |
| PACIBatch=128, Accumulation Factor (a)=82026.06 | 0.0028 | 0.0018 | |
| PACIBatch=128, Accumulation Factor (a)=42026.06 | 0.0028 | 0.0021 | |
| SoftMoE*k=1, alpha=4, Train-AE=3.64, Infer-AE=3.73, Active-Param=242.06, #Expert=≤ 42026.06 | 2.7 | — | |
| SoftMoEk=2, alpha=2, Train-AE=3.18, Infer-AE=3.50, Active-Param=230.74, #Expert=≤ 42026.06 | 2.74 | — | |
| Sparse MoEk=4, alpha=-, Train-AE=4, Infer-AE=4, Active-Param=255.34, #Expert=≤ 42026.06 | 2.75 | — | |
| SoftMoE*k=2, alpha=2, Train-AE=2.62, Infer-AE=3.12, Active-Param=212.05, #Expert=≤ 42026.06 | 2.75 | — | |
| SoftMoEk=1.5, alpha=2, Train-AE=1.63, Infer-AE=1.96, Active-Param=154.98, #Expert=≤ 22026.06 | 2.76 | — | |
| Sparse MoEk=3, alpha=-, Train-AE=3, Infer-AE=3, Active-Param=206.14, #Expert=≤ 32026.06 | 2.77 | — | |
| SoftMoE*k=1.5, alpha=2, Train-AE=1.53, Infer-AE=1.73, Active-Param=143.66, #Expert=≤ 22026.06 | 2.78 | — | |
| Sparse MoEk=2, alpha=-, Train-AE=2, Infer-AE=2, Active-Param=156.94, #Expert=≤ 22026.06 | 2.79 | — | |
| Sparse MoEk=1, alpha=-, Train-AE=1, Infer-AE=1, Active-Param=107.74, #Expert== 12026.06 | 2.84 | — |