arXiv · 1904.13360
Finite-Memory Strategies in POMDPs with Long-Run Average Objectives
Abstract
Partially observable Markov decision processes (POMDPs) are standard models for dynamic systems with probabilistic and nondeterministic behaviour in uncertain environments. We prove that in POMDPs with long-run average objective, the decision maker has approximately optimal strategies with finite memory. This implies notably that approximating the long-run value is recursively enumerable, as well as a weak continuity property of the value with respect to the transition function.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Krishnendu Chatterjee, Raimundo Saona, Bruno Ziliotto. 2019-04-30. Finite-Memory Strategies in POMDPs with Long-Run Average Objectives. https://doi.org/10.1287/moor.2020.1116
Cite the original work for its findings. Save a collection to share your selection of sources.