arXiv · 2606.21354
Restart and Adaptive Acceleration in Stochastic Gradient Methods
Abstract
We study restart schemes in stochastic optimization problems for non-smooth and weakly convex that satisfy a Kurdyka-\L ojasiewicz inequality. We show that using restarts allows us to leverage the K{\L} inequalities to achieve improved rates of convergence, with acceleration depending explicitly on the K{\L} exponent. Furthermore, optimal restart schedules lead to learning-rates akin to Polyak steps for SGD. While regularity constants such as the K{\L} exponent are typically unknown in practice, we prove that restart schemes are robust to a significant misspecification of these constants, hence nearly adaptive. We detail numerical experiments on both toy problems, where the K{\L} exponent is controlled, and training of Large Language Models (LLMs).
Explore related subjects
Keep this discovery
Ali Elhishi, Chistophe Roux, Alexandre d'Aspremont. 2026-06-19. Restart and Adaptive Acceleration in Stochastic Gradient Methods. https://arxiv.org/abs/2606.21354
Cite the original work for its findings. Save a collection to share your selection of sources.