arXiv · 2605.26000
Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance
Abstract
Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization. However, in some modern statistical learning problems, stochastic gradients can exhibit infinite-variance behavior. Consequently, classical inference methods for SGD that rely on a finite-variance assumption break down. We develop a model-agnostic methodology for constructing confidence regions from SGD iterates in both the finite- and infinite-variance regimes. We first show that Polyak--Ruppert averaging has an asymptotic directional scale no larger than that of the fastest-rate final iterate, analogous to its lower asymptotic variance in the finite-variance setting. Accordingly, we focus our inference methodology on the Polyak--Ruppert averaged estimator. Specifically, we establish a joint central limit theorem for this estimator and an empirical second-moment normalizer from the same iterates. This joint limit yields a self-normalized statistic in which the leading tail-dependent scaling terms cancel. We then use subsampling to estimate the relevant quantiles, avoiding explicit estimation of nuisance parameters including tail indices, slowly varying functions, or stable-law parameters. The resulting confidence regions are straightforward to implement and asymptotically valid in both the finite- and infinite-variance regimes. Empirical studies show reliable coverage in various settings, supporting the proposed method as a practical tool for uncertainty quantification in stochastic optimization.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jose Blanchet, Peter Glynn, Wenhao Yang. 2026-05-25. Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance. https://arxiv.org/abs/2605.26000
Cite the original work for its findings. Save a collection to share your selection of sources.