arXiv · 1507.07984
A constrained optimization perspective on actor critic algorithms and application to network routing
Abstract
We propose a novel actor-critic algorithm with guaranteed convergence to an optimal policy for a discounted reward Markov decision process. The actor incorporates a descent direction that is motivated by the solution of a certain non-linear optimization problem. We also discuss an extension to incorporate function approximation and demonstrate the practicality of our algorithms on a network routing application.
Explore related subjects
Keep this discovery
Prashanth L. A., H. L. Prasad, Shalabh Bhatnagar, Prakash Chandra. 2015-07-28. A constrained optimization perspective on actor critic algorithms and application to network routing. https://arxiv.org/abs/1507.07984
Cite the original work for its findings. Save a collection to share your selection of sources.