arXiv · 2107.04422
Policy Gradient Methods for Distortion Risk Measures
Abstract
We propose policy gradient algorithms which learn risk-sensitive policies in a reinforcement learning (RL) framework. Our proposed algorithms maximize the distortion risk measure (DRM) of the cumulative reward in an episodic Markov decision process in on-policy and off-policy RL settings, respectively. We derive a variant of the policy gradient theorem that caters to the DRM objective, and integrate it with a likelihood ratio-based gradient estimation scheme. We derive non-asymptotic bounds that establish the convergence of our proposed algorithms to an approximate stationary point of the DRM objective.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nithia Vijayan, Prashanth L. A. 2021-07-09. Policy Gradient Methods for Distortion Risk Measures. https://arxiv.org/abs/2107.04422
Cite the original work for its findings. Save a collection to share your selection of sources.