arXiv · 2003.12895
Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks
Abstract
We prove that a single step of gradient decent over depth two network, with $q$ hidden neurons, starting from orthogonal initialization, can memorize $\Omega\left(\frac{dq}{\log^4(d)}\right)$ independent and randomly labeled Gaussians in $\mathbb{R}^d$. The result is valid for a large class of activation functions, which includes the absolute value.
Explore related subjects
Keep this discovery
Amit Daniely. 2020-03-28. Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks. https://arxiv.org/abs/2003.12895
Cite the original work for its findings. Save a collection to share your selection of sources.