SearcharxivSearch

arXiv subjects

Gerard Ben Arous

Publications and source records attributed to Gerard Ben Arous.

At least 19 recordsLinked to original sources

Wandering Exponents and the Free Energy of the High-Dimensional Elastic Polymer

We study the behavior of the elastic polymer, a model of a directed polymer in a continuous Gaussian random environment that is independent in time and correlated in space, as the dimension of the environment is taken to infinity. We give an explicit asymptotic formula for the free energy, which is given in terms of the distribution of the inner product of two sampled configurations, which we also obtain an implicit formula for. From this, we provide an explicit characterization of both the low- and high-temperature phases of this model in terms of the spatial correlation function of the environment. We find asymptotics for the wandering exponent when the spatial correlation function is either an exponential or a power-law decay. Our results show that when the correlations are either suitably weak or short ranged, the model is asymptotically diffusive. On the other hand, for suitably strong long ranged correlations, the model is asymptotically superdiffusive. Moreover, we show that this transition coincides exactly with another transition where the model goes from being one-step replica symmetry breaking to full-step replica symmetry breaking. This rigorously confirms many of the findings of Mezard and Parisi [53] in the physics literature.

math.PR

Local geometry of high-dimensional mixture models: Effective spectral theory and dynamical transitions

We study the local geometry of empirical risks in high dimensions via the spectral theory of their Hessian and information matrices. We focus on settings where the data, $(Y_\ell)_{\ell =1}^n \in \mathbb{R}^d$, are i.i.d. draws of a $k$-Gaussian mixture model, and the loss depends on the projection of the data into a fixed number of vectors, namely $\mathbf{x}^\top Y$, where $\mathbf{x}\in \mathbb{R}^{d\times C}$ are the parameters, and $C$ need not equal $k$. This setting captures a broad class of problems such as classification by one and two-layer networks and regression on multi-index models. We provide exact formulas for the limits of the empirical spectral distribution and outlier eigenvalues and eigenvectors of such matrices in the proportional asymptotics limit, where the number of samples and dimension $n,d\to\infty$ and $n/d=ϕ\in (0,\infty)$. These limits depend on the parameters $\mathbf{x}$ only through the summary statistic of the $(C+k)\times (C+k)$ Gram matrix of the parameters and class means, $\mathbf{G} = (\mathbf{x},\boldsymbolμ)^\top(\mathbf{x},\boldsymbolμ)$. It is known that under general conditions, when $\mathbf{x}$ is trained by online stochastic gradient descent, the evolution of these same summary statistics along training converges to the solution of an autonomous system of ODEs, called the effective dynamics. This enables us to connect the training dynamics to the spectral theory of these matrices generated with test data. We demonstrate our general results by analyzing the effective spectrum along the effective dynamics in the case of multi-class logistic regression. In this setting, the empirical Hessian and information matrices have substantially different spectra, each with their own static and even dynamical spectral transitions.

math.ST

Spectral alignment of stochastic gradient descent for high-dimensional classification tasks

We rigorously study the relation between the training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, both the SGD trajectory and emergent outlier eigenspaces of the Hessian and gradient matrices align with a common low-dimensional subspace. Moreover, in multi-layer settings this alignment occurs per layer, with the final layer's outlier eigenspace evolving over the course of training, and exhibiting rank deficiency when the SGD converges to sub-optimal classifiers. This establishes some of the rich predictions that have arisen from extensive numerical studies in the last decade about the spectra of Hessian and information matrices over the course of training in overparametrized networks.

cs.LG

The Larkin Mass and Replica Symmetry Breaking in the Elastic Manifold

This is the second of a series of three papers about the Elastic Manifold model. This classical model proposes a rich picture due to the competition between the inherent disorder and the smoothing effect of elasticity. In this paper, we analyze our variational formula for the free energy obtained in our first companion paper [16]. We show that this variational formula may be simplified to one which is solved by a unique saddle point. We show that this saddle point may be solved for in terms of the corresponding critical point equation. Moreover, its terms may be interpreted in terms of natural statistics of the model: namely the overlap distribution and effective radius of the model at a given site. Using this characterization, obtain a complete characterization of the replica symmetry breaking phase. From this we are able to confirm a number of physical predictions about this boundary, namely those involving the Larkin mass [6, 53, 54], an important critical mass for the system. The zero-temperature Larkin mass has recently been shown to be the topological trivialization threshold, following work of Fyodorov and Le Doussal [37, 38], made rigorous by the first author, Bourgade and McKenna [12, 13].

math.PR

The Free Energy of the Elastic Manifold

This is the first of a series of three papers about the Elastic Manifold model. This classical model proposes a rich picture due to the competition between the inherent disorder and the smoothing effect of elasticity. In this paper, we prove a Parisi formula, i.e. we compute the asymptotic quenched free energy and show it is given by the solution to a certain variational problem. This work comes after a long and distinguished line of work in the Physics literature, going back to the 1980's (including the foundational work by Daniel Fisher [29], Marc Mezard and Giorgio Parisi [50, 51], and more recently by Yan Fyodorov and Pierre Le Doussal [34, 35]. Even though the mathematical study of Spin Glasses has seen deep progress in the recent years, after the celebrated work by Michel Talagrand [67, 68], the Elastic Manifold model has been studied from a mathematical perspective, only recently and at zero temperature. The annealed topological complexity has been computed, by the first author with Paul Bourgade and Benjamin McKenna [15, 16]. Here we begin the study of this model at positive temperature by computing the quenched free energy. We obtain our Parisi formula by first applying Laplace's method to reduce the question to a related new family of spherical Spin Glass models with an elastic interaction. The upper bound is then obtained through an interpolation argument initially developed by Francisco Guerra [42] for the study of Spin Glasses. The lower bound follows by adapting the cavity method along the lines explored by Wei-Kuo Chen [23] and the multi-species synchronization method of Dmitry Panchenko [55]. In our next papers [19, 20] we will analyze the consequences of this Parisi formula.

math.PR

High-dimensional limit theorems for SGD: Effective dynamics and critical scaling

We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime. We prove limit theorems for the trajectories of summary statistics (i.e., finite-dimensional functions) of SGD as the dimension goes to infinity. Our approach allows one to choose the summary statistics that are tracked, the initialization, and the step-size. It yields both ballistic (ODE) and diffusive (SDE) limits, with the limit depending dramatically on the former choices. We show a critical scaling regime for the step-size, below which the effective ballistic dynamics matches gradient flow for the population loss, but at which, a new correction term appears which changes the phase diagram. About the fixed points of this effective dynamics, the corresponding diffusive limits can be quite complex and even degenerate. We demonstrate our approach on popular examples including estimation for spiked matrix and tensor models and classification via two-layer networks for binary and XOR-type Gaussian mixture models. These examples exhibit surprising phenomena including multimodal timescales to convergence as well as convergence to sub-optimal solutions with probability bounded away from zero from random (e.g., Gaussian) initializations. At the same time, we demonstrate the benefit of overparametrization by showing that the latter probability goes to zero as the second layer width grows.

stat.ML

Sharp complexity asymptotics and topological trivialization for the (p, k) spiked tensor model

We provide O(1) asymptotics for the average number of deep minima of the (p,k) spiked tensor model. We also derive an explicit formula for the limiting ground state energy on the N-dimensional sphere, similar to the work of Jagannath-Lopatto-Miolane. Moreover, when the signal to noise ratio is large enough, the expected number of deep minima is asymptotically finite as N tends to infinity and we determine its limit as the signal-to-noise ratio diverges.

math.PR

Online stochastic gradient descent on non-convex losses from high-dimensional inference

Stochastic gradient descent (SGD) is a popular algorithm for optimization problems arising in high-dimensional inference tasks. Here one produces an estimator of an unknown parameter from independent samples of data by iteratively optimizing a loss function. This loss function is random and often non-convex. We study the performance of the simplest version of SGD, namely online SGD, from a random start in the setting where the parameter space is high-dimensional. We develop nearly sharp thresholds for the number of samples needed for consistent estimation as one varies the dimension. Our thresholds depend only on an intrinsic property of the population loss which we call the information exponent. In particular, our results do not assume uniform control on the loss itself, such as convexity or uniform derivative bounds. The thresholds we obtain are polynomial in the dimension and the precise exponent depends explicitly on the information exponent. As a consequence of our results, we find that except for the simplest tasks, almost all of the data is used simply in the initial search phase to obtain non-trivial correlation with the ground truth. Upon attaining non-trivial correlation, the descent is rapid and exhibits law of large numbers type behavior. We illustrate our approach by applying it to a wide set of inference tasks such as phase retrieval, and parameter estimation for generalized linear models, online PCA, and spiked tensor models, as well as to supervised learning for single-layer networks with general activation functions.

stat.ML

Bounding flows for spherical spin glass dynamics

We introduce a new approach to studying spherical spin glass dynamics based on differential inequalities for one-time observables. Using this approach, we obtain an approximate phase diagram for the evolution of the energy $H$ and its gradient under Langevin dynamics for spherical $p$-spin models. We then derive several consequences of this phase diagram. For example, at any temperature, uniformly over all starting points, the process must reach and remain in an absorbing region of large negative values of $H$ and large (in norm) gradients in order 1 time. Furthermore, if the process starts in a neighborhood of a critical point of $H$ with negative energy, then both the gradient and energy must increase macroscopically under this evolution, even if this critical point is a saddle with index of order $N$. As a key technical tool, we estimate Sobolev norms of spin glass Hamiltonians, which are of independent interest.

math.PR

Algorithmic thresholds for tensor PCA

We study the algorithmic thresholds for principal component analysis of Gaussian $k$-tensors with a planted rank-one spike, via Langevin dynamics and gradient descent. In order to efficiently recover the spike from natural initializations, the signal to noise ratio must diverge in the dimension. Our proof shows that the mechanism for the success/failure of recovery is the strength of the "curvature" of the spike on the maximum entropy region of the initial data. To demonstrate this, we study the dynamics on a generalized family of high-dimensional landscapes with planted signals, containing the spiked tensor models as specific instances. We identify thresholds of signal-to-noise ratios above which order 1 time recovery succeeds; in the case of the spiked tensor model these match the thresholds conjectured for algorithms such as Approximate Message Passing. Below these thresholds, where the curvature of the signal on the maximal entropy region is weak, we show that recovery from certain natural initializations takes at least stretched exponential time. Our approach combines global regularity estimates for spin glasses with point-wise estimates, to study the recovery problem by a perturbative approach.

math.PR

Complex energy landscapes in spiked-tensor and simple glassy models: ruggedness, arrangements of local minima and phase transitions

We study rough high-dimensional landscapes in which an increasingly stronger preference for a given configuration emerges. Such energy landscapes arise in glass physics and inference. In particular we focus on random Gaussian functions, and on the spiked-tensor model and generalizations. We thoroughly analyze the statistical properties of the corresponding landscapes and characterize the associated geometrical phase transitions. In order to perform our study, we develop a framework based on the Kac-Rice method that allows to compute the complexity of the landscape, i.e. the logarithm of the typical number of stationary points and their Hessian. This approach generalizes the one used to compute rigorously the annealed complexity of mean-field glass models. We discuss its advantages with respect to previous frameworks, in particular the thermodynamical replica method which is shown to lead to partially incorrect predictions.

cond-mat.dis-nn

The landscape of the spiked tensor model

We consider the problem of estimating a large rank-one tensor ${\boldsymbol u}^{\otimes k}\in({\mathbb R}^{n})^{\otimes k}$, $k\ge 3$ in Gaussian noise. Earlier work characterized a critical signal-to-noise ratio $λ_{Bayes}= O(1)$ above which an ideal estimator achieves strictly positive correlation with the unknown vector of interest. Remarkably no polynomial-time algorithm is known that achieved this goal unless $λ\ge C n^{(k-2)/4}$ and even powerful semidefinite programming relaxations appear to fail for $1\ll λ\ll n^{(k-2)/4}$. In order to elucidate this behavior, we consider the maximum likelihood estimator, which requires maximizing a degree-$k$ homogeneous polynomial over the unit sphere in $n$ dimensions. We compute the expected number of critical points and local maxima of this objective function and show that it is exponential in the dimensions $n$, and give exact formulas for the exponential growth rate. We show that (for $λ$ larger than a constant) critical points are either very close to the unknown vector ${\boldsymbol u}$, or are confined in a band of width $Θ(λ^{-1/(k-1)})$ around the maximum circle that is orthogonal to ${\boldsymbol u}$. For local maxima, this band shrinks to be of size $Θ(λ^{-1/(k-2)})$. These `uninformative' local maxima are likely to cause the failure of optimization algorithms.

math.ST

Explorations on high dimensional landscapes

Finding minima of a real valued non-convex function over a high dimensional space is a major challenge in science. We provide evidence that some such functions that are defined on high dimensional domains have a narrow band of values whose pre-image contains the bulk of its critical points. This is in contrast with the low dimensional picture in which this band is wide. Our simulations agree with the previous theoretical work on spin glasses that proves the existence of such a band when the dimension of the domain tends to infinity. Furthermore our experiments on teacher-student networks with the MNIST dataset establish a similar phenomenon in deep networks. We finally observe that both the gradient descent and the stochastic gradient descent methods can reach this level within the same number of steps.

stat.ML

Biased random walks on random graphs

These notes cover one of the topics programmed for the St Petersburg School in Probability and Statistical Physics of June 2012. The aim is to review recent mathematical developments in the field of random walks in random environment. Our main focus will be on directionally transient and reversible random walks on different types of underlying graph structures, such as $\mathbb{Z}$, trees and $\mathbb{Z}^d$ for $d\geq 2$.

math.PR

Complexity of random smooth functions on the high-dimensional sphere

We analyze the landscape of general smooth Gaussian functions on the sphere in dimension $N$, when $N$ is large. We give an explicit formula for the asymptotic complexity of the mean number of critical points of finite and diverging index at any level of energy and for the mean Euler characteristic of level sets. We then find two possible scenarios for the bottom landscape, one that has a layered structure of critical values and a strong correlation between indexes and critical values and another where even at levels below the limiting ground state energy the mean number of local minima is exponentially large. We end the paper by discussing how these results can be interpreted in the language of spin glasses models.

math.PR

A Central Limit Theorem in Many-Body Quantum Dynamics

We study the many body quantum evolution of bosonic systems in the mean field limit. The dynamics is known to be well approximated by the Hartree equation. So far, the available results have the form of a law of large numbers. In this paper we go one step further and we show that the fluctuations around the Hartree evolution satisfy a central limit theorem. Interestingly, the variance of the limiting Gaussian distribution is determined by a time-dependent Bogoliubov transformation describing the dynamics of initial coherent states in a Fock space representation of the system.

math-ph

Einstein relation for biased random walk on Galton--Watson trees

We prove the Einstein relation, relating the velocity under a small perturbation to the diffusivity in equilibrium, for certain biased random walks on Galton--Watson trees. This provides the first example where the Einstein relation is proved for motion in random media with arbitrary deep traps.

math.PR