arXiv · 2610.08534
How Bregman Divergences Shape Shampoo
Abstract
Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bing Liu, Wenjie Zhou, Chengcheng Zhao, Hongtao Zhang, Boao Kong, Felix Dangel, Wu Lin. 2026-10-06. How Bregman Divergences Shape Shampoo. https://arxiv.org/abs/2610.08534
Cite the original work for its findings. Save a collection to share your selection of sources.