Image Classification on CIFAR-10 (17 test)
90.2Accuracy (alpha=1.0)All-reduce + momentum
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| All-reduce + momentumTopology=fully connected, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=Standard momentum2021.10 | 90.2 | 90.2 | 90.2 | |
| RelaySGD + local momentumTopology=binary trees, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=Local momentum2021.10 | 90.2 | 89.5 | 89.1 | |
| DP-SGD + quasi-global momentumTopology=ring, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=Quasi-global momentum2021.10 | 89.5 | 84.8 | 63.3 | |
| Stochastic gradient push + local momentumTopology=time-varying exponential, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=Local momentum2021.10 | 89.5 | 89.2 | 87.5 | |
| D2 + local momentumTopology=ring, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=Local momentum2021.10 | 88.2 | 88.5 | 61 | |
| RelaySGDTopology=binary trees, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=None2021.10 | 87.4 | 86.9 | 84.6 | |
| DP-SGDTopology=ring, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=None2021.10 | 87.4 | 79.9 | 53.9 | |
| Stochastic gradient pushTopology=time-varying exponential, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=None2021.10 | 87.4 | 86.7 | 86.7 | |
| D2Topology=ring, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=None2021.10 | 87.2 | 84 | 38.2 | |
| All-reduce (baseline)Topology=fully connected, Backbone=VGG-11, Number of workers=16, Communication budget per iteration=2 models, Momentum strategy=None2021.10 | 87 | 87 | 87 |