Training Efficiency on Llama3-8x70B Coarse-grained
41.6MFUMCore w/ Folding
Evaluation Results
| Method | Links | |
|---|---|---|
| MCore w/ FoldingGPUs=256, Global batch size=2562025.04 | 41.6 | |
| MCoreGPUs=256, Global batch size=2562025.04 | 38.8 | |
| FSDP + EPGPUs=256, Global batch size=2562025.04 | 19.6 |