arXiv · 2609.04794
Policy Iteration for Domain Randomized Linear Quadratic Systems
Abstract
In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability across all sampled systems at each iteration. We show that the method yields monotonic improvement of the sample-average objective and that a stabilizing step size always exists. Under standard smoothness assumptions, the iterates converge subsequentially to stationary points, and under a gradient-dominance condition, we obtain a global linear convergence rate.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abbas Pasdar, Farnaz Adib Yaghmaie. 2026-09-04. Policy Iteration for Domain Randomized Linear Quadratic Systems. https://arxiv.org/abs/2609.04794
Cite the original work for its findings. Save a collection to share your selection of sources.