arXiv · 2605.12994
DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum
Abstract
We study differentially private optimization with matrix-orthogonalized momentum. DP-Muon uses conventional global per-example clipping and one Gaussian gradient release per step; matrix updates and auxiliary updates are post-processing. Our main contribution concerns the additional mean distortion created when fresh Gaussian noise passes through a nonlinear matrix map. Conditioning on the actual adaptive history immediately before the current noise yields an exact Gaussian heat identity. For a smooth Newton-Schulz map, first-order DP-MuonBC reduces this conditional output bias from second to fourth order in the fresh noise scale, and an arbitrary-order extension has bias of order $2K+2$. We prove matrix-block stationarity bounds under global clipping, retain finite-step orthogonalization error explicitly, and give an exact criterion for improvement of the resulting upper bound. A separate inequality exposes the effect of auxiliary Adam updates. GPT-2 experiments on E2E at four privacy targets favor the reported Muon configurations over Adam baselines in test NLL.
Explore related subjects
Keep this discovery
Jihwan Kim, Chenglin Fan. 2026-05-13. DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum. https://arxiv.org/abs/2605.12994
Cite the original work for its findings. Save a collection to share your selection of sources.