arXiv · 2605.07184
Convergence of Stochastic Gradient Descent with mini-batching and infinite variance
Abstract
Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing batch sizes when the gradient noise belongs to the domain of attraction of an $\alpha$-stable law with $\alpha\in(1,2)$. Building on existing results for the finite-variance regime and for heavy-tailed SGD without batching, we establish three main results. First, we derive $L^p$ moment bounds for the SGD error and show that increasing batch sizes lead to faster convergence rates. In particular, batching enables convergence in probability even for a constant stepsize. Second, we prove that the properly normalized SGD iterates converge in distribution to the stationary law of an Ornstein-Uhlenbeck process driven by an $\alpha$-stable L\'evy process. Third, for Polyak-Ruppert averaging we obtain a stable limit theorem with a normalization that explicitly depends on the batch-size schedule.
Explore related subjects
Keep this discovery
Bartosz Glowacki, Rafal Kulik, Philippe Soulier. 2026-05-08. Convergence of Stochastic Gradient Descent with mini-batching and infinite variance. https://arxiv.org/abs/2605.07184
Cite the original work for its findings. Save a collection to share your selection of sources.