Training Efficiency on Mixtral-8x22B Coarse-grained
49.3MFUMCore w/ Folding
Evaluation Results
| Method | Links | |
|---|---|---|
| MCore w/ FoldingGPUs=128, Global batch size=2562025.04 | 49.3 | |
| MCoreGPUs=128, Global batch size=2562025.04 | 46.3 | |
| TP+EP+DPGPUs=128, Global batch size=2562025.04 | 36.6 | |
| FSDP + EPGPUs=128, Global batch size=2562025.04 | 23.4 | |
| FSDPGPUs=128, Global batch size=2562025.04 | 4.3 |