SearcharxivSearch

arXiv subjects

Anil Kamber

Publications and source records attributed to Anil Kamber.

3 recordsLinked to original sources

On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness

Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding of its effect on the loss landscape and the set of minimizers remains limited. In this paper, we show that $\ell^2$-regularized deep matrix factorization/deep linear network training problems with squared-error loss admit a unique end-to-end minimizer for all target matrices subject to factorization, except for a set of Lebesgue measure zero formed by the depth and the regularization parameter. This observation reveals fundamental properties of the loss landscape of regularized deep matrix factorization problems: the Hessian spectrum is constant across all minimizers of the regularized deep scalar factorization problem with squared-error loss. Moreover, we show that, in regularized deep matrix factorization problems with squared-error loss, if the target matrix does not belong to the Lebesgue measure-zero set, then the Frobenius norm of each layer is constant across all minimizers. This, in turn, yields a global lower bound on the trace of the Hessian evaluated at any minimizer of the regularized deep matrix factorization problem. Furthermore, we establish a critical threshold for the regularization parameter above which the unique end-to-end minimizer collapses to zero.

stat.ML

Sharpness of Minima in Deep Matrix Factorization

Understanding the geometry of the loss landscape near a minimum is key to explaining the implicit bias of gradient-based methods in non-convex optimization problems such as deep neural network training and deep matrix factorization. A central quantity to characterize this geometry is the maximum eigenvalue of the Hessian of the loss. Currently, its precise role has been obfuscated because no exact expressions for this sharpness measure were known in general settings. In this paper, we present the first exact expression for the maximum eigenvalue of the Hessian of the squared-error loss at any minimizer in deep matrix factorization/deep linear neural network training problems, resolving an open question posed by Mulayoff & Michaeli (2020). This expression reveals a fundamental property of the loss landscape in deep matrix factorization: Having a constant product of the spectral norms of the left and right intermediate factors across layers is a sufficient condition for flatness. Most notably, in both depth-$2$ matrix and deep overparameterized scalar factorization, we show that this condition is both necessary and sufficient for flatness, which implies that flat minima are spectral-norm balanced even though they are not necessarily Frobenius-norm balanced. To complement our theory, we provide the first empirical characterization of an escape phenomenon during gradient-based training near a minimizer of a deep matrix factorization problem.

stat.ML

Half-Space Modeling with Reflecting Surface in Molecular Communication

In Molecular Communications via Diffusion (MCvD), messenger molecules are emitted by a transmitter and propagate randomly through the fluidic environment. In biological systems, the environment can be considered a bounded space, surrounded by various structures such as tissues and organs. The propagation of molecules is affected by these structures, which reflect the molecules upon collision. Deriving the channel response of MCvD systems with an absorbing spherical receiver requires solving the 3-D diffusion equation in the presence of reflecting and absorbing boundary conditions, which is extremely challenging. In this paper, the method of images is brought to molecular communication (MC) realm to find a closed-form solution to the channel response of a single-input single-output (SISO) system near an infinite reflecting surface. We showed that a molecular SISO system in a 3-D half-space with an infinite reflecting surface could be approximated as a molecular single-input multiple-output (SIMO) system in a 3-D space, which consists of two symmetrically located, with respect to the reflecting surface, identical absorbing spherical receivers.

cs.IT