arXiv · 2105.04218
Exploiting Elasticity in Tensor Ranks for Compressing Neural Networks
Abstract
Elasticities in depth, width, kernel size and resolution have been explored in compressing deep neural networks (DNNs). Recognizing that the kernels in a convolutional neural network (CNN) are 4-way tensors, we further exploit a new elasticity dimension along the input-output channels. Specifically, a novel nuclear-norm rank minimization factorization (NRMF) approach is proposed to dynamically and globally search for the reduced tensor ranks during training. Correlation between tensor ranks across multiple layers is revealed, and a graceful tradeoff between model size and accuracy is obtained. Experiments then show the superiority of NRMF over the previous non-elastic variational Bayesian matrix factorization (VBMF) scheme.
Explore related subjects
Keep this discovery
Jie Ran, Rui Lin, Hayden K. H. So, Graziano Chesi, Ngai Wong. 2021-05-10. Exploiting Elasticity in Tensor Ranks for Compressing Neural Networks. https://arxiv.org/abs/2105.04218
Cite the original work for its findings. Save a collection to share your selection of sources.