arXiv · 1910.08446
Autonomous exploration for navigating in non-stationary CMPs
Abstract
We consider a setting in which the objective is to learn to navigate in a controlled Markov process (CMP) where transition probabilities may abruptly change. For this setting, we propose a performance measure called exploration steps which counts the time steps at which the learner lacks sufficient knowledge to navigate its environment efficiently. We devise a learning meta-algorithm, MNM and prove an upper bound on the exploration steps in terms of the number of changes.
Explore related subjects
Keep this discovery
Pratik Gajane, Ronald Ortner, Peter Auer, Csaba Szepesvari. 2019-10-18. Autonomous exploration for navigating in non-stationary CMPs. https://arxiv.org/abs/1910.08446
Cite the original work for its findings. Save a collection to share your selection of sources.