arXiv · 2607.12590
Environment Parameter Gradient Theorem for Co-Design in Reinforcement Learning
Abstract
Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable, i.e., physical or operational parameters can be tuned to shape the system's transition dynamics and costs experienced by the RL agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem --- a formal expression for the gradient of the RL's objective function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{\pi,\xi}(s,a,\zeta)$, which comprises two copies of the environment parameters: $\zeta$ governs the cost and transition dynamics at the current state--action pair, while $\xi$ governs the future rollouts. This decoupling yields a tractable closed-form gradient expression and is essential to the theorem's derivation. Building on this result, we develop a model-free algorithm that simultaneously learns the optimal policy and the environment parameters. We demonstrate the efficacy of our framework on a UAV network design problem, where the optimal UAV placement (environment parameters) and communication routes (governed by the policy) are learned jointly to minimize the total communication cost in the network.
Explore related subjects
Keep this discovery
Amber Srivastava. 2026-07-14. Environment Parameter Gradient Theorem for Co-Design in Reinforcement Learning. https://arxiv.org/abs/2607.12590
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.