arXiv · 1611.07476
Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
Abstract
We look at the eigenvalues of the Hessian of a loss function before and after training. The eigenvalue distribution is seen to be composed of two parts, the bulk which is concentrated around zero, and the edges which are scattered away from zero. We present empirical evidence for the bulk indicating how over-parametrized the system is, and for the edges that depend on the input data.
Explore related subjects
Keep this discovery
Levent Sagun, Leon Bottou, Yann LeCun. 2016-11-22. Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond. https://arxiv.org/abs/1611.07476
Cite the original work for its findings. Save a collection to share your selection of sources.