Training Efficiency on Qwen2-57B-A14B Fine-grained
39MFUMCore w/ Folding
Evaluation Results
| Method | Links | |
|---|---|---|
| MCore w/ FoldingGPUs=64, Global batch size=2562025.04 | 39 | |
| MCoreGPUs=64, Global batch size=2562025.04 | 35.3 | |
| FSDP + EPGPUs=64, Global batch size=2562025.04 | 25.4 | |
| TP+EP+DPGPUs=64, Global batch size=2562025.04 | 23.1 | |
| FSDPGPUs=64, Global batch size=2562025.04 | 9.9 |