arXiv · 2203.09044
Convert, compress, correct: Three steps toward communication-efficient DNN training
Abstract
In this paper, we introduce a novel algorithm, $\mathsf{CO}_3$, for communication-efficiency distributed Deep Neural Network (DNN) training. $\mathsf{CO}_3$ is a joint training/communication protocol, which encompasses three processing steps for the network gradients: (i) quantization through floating-point conversion, (ii) lossless compression, and (iii) error correction. These three components are crucial in the implementation of distributed DNN training over rate-constrained links. The interplay of these three steps in processing the DNN gradients is carefully balanced to yield a robust and high-performance scheme. The performance of the proposed scheme is investigated through numerical evaluations over CIFAR-10.
Explore related subjects
Keep this discovery
Zhong-Jing Chen, Eduin E. Hernandez, Yu-Chih Huang, Stefano Rini. 2022-03-17. Convert, compress, correct: Three steps toward communication-efficient DNN training. https://arxiv.org/abs/2203.09044
Cite the original work for its findings. Save a collection to share your selection of sources.