arXiv · 1706.06643
Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines
Abstract
We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton et al. 2000).
Explore related subjects
Keep this discovery
Philip S. Thomas, Emma Brunskill. 2017-06-20. Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines. https://arxiv.org/abs/1706.06643
Cite the original work for its findings. Save a collection to share your selection of sources.