arXiv · 2207.12505
On the benefits of non-linear weight updates
Abstract
Recent work has suggested that the generalisation performance of a DNN is related to the extent to which the Signal-to-Noise Ratio is optimised at each of the nodes. In contrast, Gradient Descent methods do not always lead to SNR-optimal weight configurations. One way to improve SNR performance is to suppress large weight updates and amplify small weight updates. Such balancing is already implicit in some common optimizers, but we propose an approach that makes this explicit. The method applies a non-linear function to gradients prior to making DNN parameter updates. We investigate the performance with such non-linear approaches. The result is an adaptation to existing optimizers that improves performance for many problem types.
Explore related subjects
Keep this discovery
Paul Norridge. 2022-07-25. On the benefits of non-linear weight updates. https://arxiv.org/abs/2207.12505
Cite the original work for its findings. Save a collection to share your selection of sources.