SearcharxivSearch

arXiv · 2306.13326

Solving systems of Random Equations via First and Second-Order Optimization Algorithms

Abstract

We revisit the problem of solving $n$ random equations in $d$ real variables, when the equations are independent realizations of a Gaussian process in $d$ dimensions. A special case is the one of random polynomial equations, which has been studied since Littlewood-Offord and Kac in the 1940s (who studied of existence of solutions of random polynomials) and Shub and Smale in the 1990s. The last authors first investigated the computational aspect of this problem. Smale's `17th problem' asks whether a system of random polynomial equations can be (approximately) solved in average case polynomial time. We formulate this as a nonconvex optimization problem, and apply local algorithms based on gradient or Hessian information. We leverage recent advances in spin glass theory to characterize the optimal algorithm in this class, and show that the latter undergoes a phase transition at a critical value $\alpha_{\text{alg}}$ of the ratio $\alpha=n/d$. We establish that near-solutions can be found with-high probability for $\alpha<\alpha_{\text{alg}}$, while a companion paper proves that a broad class of efficient algorithms fail for $\alpha>\alpha_{\text{alg}}$ (we outline the proof of this hardness result). We further prove that there are cases such that for $(1+\delta)\alpha_{\text{alg}} 0$ arbitrarily small) solutions exists with high probability but are not found efficiently by a broad class of algorithms. We compare our predictions with numerical simulations using the optimal algorithm we propose as well as stochastic gradient descent, and show that they are accurate for a related albeit non-Gaussian cost function. We finally observe empirically a sensitivity cross-over in the behavior of optimization algorithms, below $\alpha_{\text{alg}}$. This marks a qualitative departure with respect to standard optimization theories.

Explore related subjects

Keep this discovery

BibTeXRIS

Andrea Montanari, Eliran Subag. 2023-06-23. Solving systems of Random Equations via First and Second-Order Optimization Algorithms. https://arxiv.org/abs/2306.13326

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Averaging principles for nonautonomous multiscale stochastic Burgers equations with reflection

In this paper, we study averaging principles for nonautonomous multiscale stochastic Burgers equations with reflection. First, we derive a general averaging principle applicable to such equations under minimal assumptions. Subsequently, since the coefficients of the obtained averaged equation still depend on the small scaling parameter $\e$, we impose either periodic or asymptotic conditions on the coefficients, thereby obtain two distinct averaged equations whose coefficients are independent of $\e$ and establish two averaging principles. Stopping times and Khasminskii's time discretization schemes play an important role. Finally, a concrete example is provided to illustrate the applicability and validity of the theoretical results.

math.PR

Spectral properties of Random Matrices

We give the theoretical foundations of random matrix theory through the definitions of a random matrix, a random probability measure and the corresponding empirical spectral distribution. The technical tool we use is the Stieltjes transform method through which we prove optimal convergence of the empirical spectral distribution of random sample covariance matrices to the deterministic Marchenko-Pastur distribution. We also give new results about the rigidity of the eigenvalues of this random sample covariance matrix and the rate of their convergence. We then define the Dyson equation method to prove new local laws about a random matrix model that interpolates between the Marchenko-Pastur distribution, the elliptical law and the circular law. Through our work these local laws can be considered universal.

math.PR

Moments approach for the elephant random walk

We discuss the method of moments for the one-dimensional elephant random walk (ERW). We first derive a differential recurrence relation for the characteristic function of the ERW, which yields a corresponding system of recurrence relations for its moments. We then obtain asymptotic approximations for the moments in each of the three parameter regimes of the ERW. Finally, by establishing the convergence of the moments and verifying the corresponding moment-determinacy conditions, we identify the limiting distributions of the ERW in each regime.

math.PR