Searcharxiv⌕ Search

arXiv · 2610.10527

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

Abstract

Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ($\mathtt{DSGD}$). For smooth non-convex costs under bounded $p$-th moment noise, $p \in (1,2]$, we show that clipped $\mathtt{DSGD}$ achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized $\mathtt{DSGD}$ can fail to converge, clipping retains magnitude information, enabling $\mathtt{DSGD}$ to be convergent and order-optimal. Numerical experiments validate our theory.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed. 2026-10-07. Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping. https://arxiv.org/abs/2610.10527

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Combining additivity and active subspaces for high-dimensional Gaussian process modeling

Gaussian processes are a widely embraced technique for regression due to their good prediction accuracy, analytical tractability, and built-in capabilities for uncertainty quantification. However, they suffer from the curse of dimensionality whenever the number of variables increases. This challenge is generally addressed by assuming additional structure in the problem, the preferred options being either additivity or low intrinsic dimensionality. After discussing their relative merits, our contribution for high-dimensional Gaussian process modeling is to combine them with a multi-fidelity strategy. We detail the corresponding construction and showcase the advantages through experiments on synthetic functions and datasets.

math.OC↗

Constrained portfolio game with heterogeneous agents

We investigate stochastic utility maximization games under relative performance concerns in both finite-agent and infinite-agent (graphon) settings. An incomplete market model is considered where agents with power (CRRA) utility functions trade in a common risk-free bond and individual stocks driven by both common and idiosyncratic noise. The Nash equilibrium for both settings is characterized by forward-backward stochastic differential equations (FBSDEs) with a quadratic growth generator, where the solution of the graphon game leads to a novel form of infinite-dimensional McKean-Vlasov FBSDEs. Under mild conditions, we prove the existence of Nash equilibrium for both the graphon game and the $n$-agent game without common noise. Furthermore, we establish a convergence result showing that, with modest assumptions on the sensitivity matrix, as the number of agents increases, the Nash equilibrium and associated equilibrium value of the finite-agent game converge to those of the graphon game.

math.OC↗

Controllability of Forward Stochastic Reaction--Convection--Diffusion Systems with Cascade Structure

We investigate the null and approximate controllability of coupled linear forward stochastic reaction--convection--diffusion systems under suitable cascade coupling conditions. The model consists of two forward stochastic parabolic equations governed by general second-order differential operators with time-, space-, and random-dependent coefficients. We consider a localized control acting on the drift term of the first equation, together with controls acting on the diffusion terms. By means of a duality argument, the controllability problem is reduced to an observability problem for the associated adjoint backward stochastic parabolic system. We establish a new global Carleman estimate for coupled backward stochastic parabolic systems whose drift terms belong to a negative Sobolev space. As an application of this estimate, we derive the required observability properties and establish the null and approximate controllability of the original forward system.

math.OC↗