Searcharxiv⌕ Search

arXiv · 2609.36600

Second-Moment Stochastic Approximation Methods

Abstract

Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tao Jiang, Lin Xiao. 2026-09-29. Second-Moment Stochastic Approximation Methods. https://arxiv.org/abs/2609.36600

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Single-loop Stochastic Riemannian ADMM for Nonsmooth Composite Optimization

We study a class of nonsmooth composite stochastic optimization on Riemannian manifolds, where the objective is the sum of an expectation function and a nonsmooth regularizer. These types of problems appear widely in various application fields, such as machine learning. Although operator-splitting methods naturally exploit this separable structure, existing Riemannian variants primarily target deterministic problems, and stochastic extensions are limited to more restrictive settings. In this work, we propose a momentum-based adaptive Riemannian stochastic alternating direction method of multipliers (MARS-ADMM), which combines a recursive variance-reduced estimator for the expectation component with a prediction-correction update for the nonsmooth block. This yields a single-loop algorithm requiring only two proximal updates, one Riemannian gradient evaluation, and $\mathcal{O}(1)$ stochastic gradient samples per iteration. Under standard assumptions, we prove that MARS-ADMM attains an $ε$-KKT point with an oracle complexity of $\mathcal{O}(ε^{-3})$, improving upon the previously best-known rate of $\widetilde{\mathcal{O}}(ε^{-3.5})$ for stochastic Riemannian primal-dual methods. This complexity also matches the best-known bounds in deterministic nonsmooth Riemannian optimization, demonstrating that deterministic-level accuracy can be achieved using only constant-size stochastic samples. Numerical experiments on two types of test problems reveal promising performances of the proposed algorithm. To the best of our knowledge, MARS-ADMM is the first stochastic Riemannian ADMM with provable optimal complexity guarantees.

math.OC↗

Convergence analysis of dynamical systems for optimization by an improved Lyapunov framework

We study the convergence analysis of continuous-time dynamical systems associated with optimization methods for strongly convex functions. Recent works have proposed systematic constructions of Lyapunov functions for such analysis, while also revealing limitations of the Lyapunov analysis. Aujol--Dossal--Rondepierre (2023) have proposed a technique to address this issue by reorganizing Lyapunov functions so as to evaluate a quantity $f(x(t)) - f_* - g(t)\|x(t)-x_*\|^2$ rather than $f(x(t)) - f_*$. By combining this technique with our computer-assisted framework to discover Lyapunov functions, we develop an improved method that reproduces an existing convergence rate or yields better rates than previous studies.

math.OC↗

Mean Field Games and Control on Large Expander Graphs

This paper investigates mean field games on sparse networks. In the case of large expander graphs, the limit topologies are analyzed using the graphexon framework, which characterizes sparse connections. We prove that the associated sequence of discrete averaging operators converges strongly to a continuous operator and this is illustrated in the development of infinite limits of Gabber-Galil-Margulis expander graphs \cite{gabber1981explicit}. These properties enable the formulation and existence proof of equilibria for linear-quadratic mean field games in which each agent is identified by a spatial network label $α\in X$ and only interacts with the neighborhood average characterized by the operator $\mathcal{G}$, i.e., the average state of a limited number of connected neighbors in a large expander graph. Furthermore, algebraic conditions induced by the spectral gap of $\mathcal{G}$ for the global asymptotic stability of the closed-loop system are established, and parameter thresholds that give rise to a Turing-type topological instability are established.

math.OC↗