arXiv · 2609.08823
Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games
Abstract
Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Finite-horizon zero-sum networked separable Markov games occupy an important middle ground: they retain global competition structure through pairwise interactions, while preserving computational tractability of Nash equilibria (NE) in the full-information and known-transition setting. Existing algorithms for this class either proceed through equilibrium-collapse arguments for a simplified setting where a single controller determines the transition probability, or backward dynamic programming that relies on equilibrium solvers at each stage. However, the design and analysis of direct policy-update approaches remain inadequate. To address this issue, we propose the entropy-regularized optimistic multiplicative weights update (ER-OMWU), a complementary single-loop policy dynamic that updates players' policies symmetrically and returns an approximate NE in the last iteration. We provide a first last-iterate convergence analysis of policy dynamics in the games of interest: after $\widetilde{O}(1/{\epsilon})$ iterations, the returned policy is an $\epsilon$-approximate Nash equilibrium. The result preserves the near-linear convergence rate achieved by policy optimization in two-player zero-sum Markov games, but extends the policy-dynamics viewpoint to a more complicated but structured multi-player setting.
Explore related subjects
Keep this discovery
Zailin Ma. 2026-09-08. Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games. https://arxiv.org/abs/2609.08823
Cite the original work for its findings. Save a collection to share your selection of sources.