arXiv · 1507.04635
Black-Box Policy Search with Probabilistic Programs
Abstract
In this work, we explore how probabilistic programs can be used to represent policies in sequential decision problems. In this formulation, a probabilistic program is a black-box stochastic simulator for both the problem domain and the agent. We relate classic policy gradient techniques to recently introduced black-box variational methods which generalize to probabilistic program inference. We present case studies in the Canadian traveler problem, Rock Sample, and a benchmark for optimal diagnosis inspired by Guess Who. Each study illustrates how programs can efficiently represent policies using moderate numbers of parameters.
Explore related subjects
Keep this discovery
Jan-Willem van de Meent, Brooks Paige, David Tolpin, Frank Wood. 2015-07-16. Black-Box Policy Search with Probabilistic Programs. https://arxiv.org/abs/1507.04635
Cite the original work for its findings. Save a collection to share your selection of sources.