TY - RPRT TI - Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning AU - Hao Yu AU - Sen Yang AU - Shenghuo Zhu PY - 2018 UR - https://arxiv.org/abs/1807.06629 ID - 1807.06629 ER -