arXiv · 2106.15594
Limited depth bandit-based strategy for Monte Carlo planning in continuous action spaces
Abstract
This paper addresses the problem of optimal control using search trees. We start by considering multi-armed bandit problems with continuous action spaces and propose LD-HOO, a limited depth variant of the hierarchical optimistic optimization (HOO) algorithm. We provide a regret analysis for LD-HOO and show that, asymptotically, our algorithm exhibits the same cumulative regret as the original HOO while being faster and more memory efficient. We then propose a Monte Carlo tree search algorithm based on LD-HOO for optimal control problems and illustrate the resulting approach's application in several optimal control problems.
Explore related subjects
Keep this discovery
Ricardo Quinteiro, Francisco S. Melo, Pedro A. Santos. 2021-06-29. Limited depth bandit-based strategy for Monte Carlo planning in continuous action spaces. https://arxiv.org/abs/2106.15594
Cite the original work for its findings. Save a collection to share your selection of sources.