arXiv · 2311.07411
A Large Deviations Perspective on Policy Gradient Algorithms
Abstract
Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly non-convex objectives satisfying a Polyak-{\L}ojasiewicz condition. Leveraging the contraction principle from large deviations theory, we illustrate the potential of this result by showing how convergence properties of policy gradient with a softmax parametrization and an entropy regularized objective can be naturally extended to a wide spectrum of other policy parametrizations.
Explore related subjects
Keep this discovery
Wouter Jongeneel, Daniel Kuhn, Mengmeng Li. 2023-11-13. A Large Deviations Perspective on Policy Gradient Algorithms. https://arxiv.org/abs/2311.07411
Cite the original work for its findings. Save a collection to share your selection of sources.