arXiv · 2609.01099
The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives
Abstract
We study partially observable Markov decision processes (POMDPs) with long-run average objectives, where the payoff is defined as the limit inferior of the expected average rewards. We consider the computational problem of approximating the long-run average value of a POMDP. In general, the long-run average value of a POMDP is neither computable nor approximable. We therefore consider the subclass of revealing POMDPs. Informally, a POMDP is revealing when the controller observes the underlying state with positive probability at each stage of the process. Our main contributions are threefold. First, we illustrate the practical relevance of this class of POMDPs through an application in control and optimization. Second, we present an exponential-time algorithm for the value approximation problem. Third, we establish EXPTIME-hardness by a reduction from the problem of almost-sure safety in POMDPs. Together, these results show that the problem of approximating the long-run average value in revealing POMDPs is EXPTIME-complete.
Explore related subjects
Keep this discovery
Ali Asadi, Krishnendu Chatterjee, David Lurie. 2026-09-01. The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives. https://arxiv.org/abs/2609.01099
Cite the original work for its findings. Save a collection to share your selection of sources.