arXiv · 2608.15824
Second-Moment Memory in Coordinatewise Adam
Abstract
Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood. We show that second-moment memory can itself suppress progress toward the optimum even under finite-variance stochastic gradients. For a simple two-point oracle, the expected positive normalized update is $O(M_2^{-1/2})$ after an initialization transient, where $M_2=(1-\beta_2)^{-1}$ is the second-moment memory length. We convert this directional bound, under the stated memory and stepsize scaling, into an average-stationarity lower bound of the same order on a smooth convex problem with normalized gap, smoothness, and variance. Long second-moment memory can slow optimization even when the gradient noise has finite variance.
Explore related subjects
Keep this discovery
Jeonseong Kim. 2026-08-16. Second-Moment Memory in Coordinatewise Adam. https://arxiv.org/abs/2608.15824
Cite the original work for its findings. Save a collection to share your selection of sources.