arXiv · 2410.02605
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
Abstract
We derive a policy gradient theorem for Cumulative Prospect Theory (CPT) objectives in finite-horizon Reinforcement Learning (RL), generalizing the standard policy gradient theorem and encompassing distortion-based risk objectives as special cases. Motivated by behavioral economics, CPT combines an asymmetric utility transformation around a reference point with probability distortion. Building on our theorem, we design a first-order policy gradient algorithm for CPT-RL using a Monte Carlo gradient estimator based on order statistics. We establish statistical guarantees for the estimator and prove asymptotic convergence of the resulting algorithm to first-order stationary points of the (generally nonconvex) CPT objective. We complement our asymptotic analysis with a non-asymptotic total sample complexity analysis to reach an approximate first-order stationary policy. Simulations illustrate qualitative behaviors induced by CPT and compare our first-order approach to existing zeroth-order methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Olivier Lepel, Anas Barakat. 2024-10-03. Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning. https://arxiv.org/abs/2410.02605
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.