Searcharxiv⌕ Search

arXiv · 2610.09912

Two Answers on Heavy-Ball Dynamics

Abstract

We answer two questions about the heavy-ball method on smooth strongly convex functions. The first concerns Polyak's tuning. Its dimension-uniform worst-case exponential rate equals the largest growth factor of an interpolable mixture of rotating geometric sequences, and we determine this factor for every condition number. The optimal mixtures form four algebraic families; the classification combines exact polynomial certificates, a rational parametrization, an exact solution in a quintic number field at the critical three-cycle, and verified interval continuation. Two frequencies suffice, one function in dimension five attains the rate at every horizon, the upper and lower bounds differ by constant factors, and the method converges on the whole class if and only if $κ<9+4\sqrt5$, the threshold located by Badithela and Seiler. The second question concerns arbitrary step sizes and momenta, in particular the regions of the parameter plane between the computed Lyapunov and cycle regions. We describe the phase diagram through the angle $Θ(ω)$ between the characteristic polynomials of the two extreme quadratics. The method converges when $Θ(ω)>ω/2$, has a $k$-periodic orbit when $Θ(2πj/k)\leπ/k$, and the remaining regions are rotation-number gaps organized by Farey fractions. A single strict certificate at one growth exponent bounds the rate, so convergence can be certified at the threshold exponent. In the first level of the gaps, fans of Farey fractions accumulating at a resonant endpoint carry periodic orbits with several harmonics, explicit fan-lag inequalities certify convergence, and both behaviors occur in the regions left open before.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dmitry Pasechnyuk-Vilensky. 2026-10-07. Two Answers on Heavy-Ball Dynamics. https://arxiv.org/abs/2610.09912

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes

Motivated by neural network training in finite-precision arithmetic environments, this work studies the convergence of perturbed iterate SGD using adaptive step sizes in an environment with numerical error. Considering a general stochastic Lipschitz continuous loss function, an asymptotic convergence result to a Clarke stationary point is proven as well as the non-asymptotic convergence to an approximate stationary point in expectation. It is assumed that only an approximation of the loss function's stochastic gradient can be computed, in addition to error in computing the SGD step itself.

math.OC↗

Sharp bounds in perturbed smooth optimization

This paper studies the problem of perturbed convex and smooth optimization. The main results describe how the solution and the value of the problem change if the objective function is perturbed. Examples include linear, quadratic, and smooth additive perturbations. Such problems naturally arise in statistics and machine learning, stochastic optimization, stability and robustness analysis, inverse problems, optimal control, etc. The results provide accurate expansions for the difference between the solution of the original problem and its perturbed counterpart with an explicit error term.

math.OC↗

Controllability Allocation Scores for Targeted Network Intervention

We introduce the controllability allocation score (CAS), a framework for determining how intervention intensity should be distributed among prescribed candidate input directions with respect to designated target variables, together with the target controllability score (TCS) as its nodewise specialization. We establish existence of the CASs and develop a general uniqueness theory based on restricted injectivity of the allocation-to-Gramian map, including generic uniqueness with respect to the time horizon. We show that restricting attention to target variables can fundamentally alter the optimal intervention allocation compared with the standard full-state setting. To enable scalability, we develop a general surrogate-Gramian framework and derive objective-performance guarantees from relative Gramian errors without requiring uniqueness or closeness of the CASs. For the TCS specialization, we further construct a target-only reduced virtual system and derive explicit bounds showing how the approximation error depends on the coupling between target and non-target nodes and on the dynamical growth rate. Experiments on human brain networks show that the reduced formulation accurately approximates the TCS at short horizons, whereas the two controllability criteria exhibit markedly different approximation accuracy at long horizons.

math.OC↗