Training Efficiency on Mixtral-8x22b-G8T8 Fine-grained
28.8MFUMCore w/ Folding
Evaluation Results
| Method | Links | |
|---|---|---|
| MCore w/ FoldingGPUs=128, Global batch size=2562025.04 | 28.8 | |
| MCoreGPUs=128, Global batch size=2562025.04 | 17.1 | |
| FSDP + EPGPUs=128, Global batch size=2562025.04 | 9 | |
| TP+EP+DPGPUs=128, Global batch size=2562025.04 | 8.7 | |
| FSDPGPUs=128, Global batch size=2562025.04 | 2.2 |