SearcharxivSearch

arXiv subjects

David Lurie

Publications and source records attributed to David Lurie.

4 recordsLinked to original sources

The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives

We study partially observable Markov decision processes (POMDPs) with long-run average objectives, where the payoff is defined as the limit inferior of the expected average rewards. We consider the computational problem of approximating the long-run average value of a POMDP. In general, the long-run average value of a POMDP is neither computable nor approximable. We therefore consider the subclass of revealing POMDPs. Informally, a POMDP is revealing when the controller observes the underlying state with positive probability at each stage of the process. Our main contributions are threefold. First, we illustrate the practical relevance of this class of POMDPs through an application in control and optimization. Second, we present an exponential-time algorithm for the value approximation problem. Third, we establish EXPTIME-hardness by a reduction from the problem of almost-sure safety in POMDPs. Together, these results show that the problem of approximating the long-run average value in revealing POMDPs is EXPTIME-complete.

math.OC

Approximating the Uniform Value in Hidden Stochastic Games with Doeblin Condition

We study \emph{zero-sum two-player hidden stochastic games}, where players receive partial observations of the state. We focus on a central solution concept for analyzing long-duration stochastic games: the \emph{uniform value}, a limiting average payoff that both players can guarantee for sufficiently long durations. In the general case, prior work provides examples of games that do not have a uniform value. Moreover, for the subclass of games that do have a uniform value, there exists no algorithm that approximates it. Therefore, we generalize the \emph{Doeblin condition} for Markov chains (which guarantees the existence of a unique invariant measure) to hidden stochastic games. Informally, the Doeblin condition for hidden stochastic games requires that, for every way to play the game, there exists a fixed belief such that, no matter the initial belief over the state of the game, after sufficiently many stages, the posterior belief is probably close to this fixed belief. Under the Doeblin condition, we prove the existence of the uniform value, provide an algorithm to approximate it, and prove that no algorithm can compute it exactly. Then, we identify structural conditions on the transition function that ensure the Doeblin condition holds both in the blind setting, where observations are uninformative, and in the hidden setting, where observations are partially informative. When considering games with only one player, namely partially observable Markov decision processes, our results provide a novel subclass in which the uniform value exists and can be approximated, but cannot be computed exactly

math.OC

Revealing POMDPs: Qualitative and Quantitative Analysis for Parity Objectives

Partially observable Markov decision processes (POMDPs) are a central model for uncertainty in sequential decision making. The most basic objective is the reachability objective, where a target set must be eventually visited, and the more general parity objectives can model all omega-regular specifications. For such objectives, the computational analysis problems are the following: (a) qualitative analysis that asks whether the objective can be satisfied with probability 1 (almost-sure winning) or probability arbitrarily close to 1 (limit-sure winning); and (b) quantitative analysis that asks for the approximation of the optimal probability of satisfying the objective. For general POMDPs, almost-sure analysis for reachability objectives is EXPTIME-complete, but limit-sure and quantitative analyses for reachability objectives are undecidable; almost-sure, limit-sure, and quantitative analyses for parity objectives are all undecidable. A special class of POMDPs, called revealing POMDPs, has been studied recently in several works, and for this subclass the almost-sure analysis for parity objectives was shown to be EXPTIME-complete. In this work, we show that for revealing POMDPs the limit-sure analysis for parity objectives is EXPTIME-complete, and even the quantitative analysis for parity objectives can be achieved in EXPTIME.

cs.CC

Uniform Value and Decidability in Ergodic Blind Stochastic Games

We study a class of two-player zero-sum stochastic games known as \textit{blind stochastic games}, where players neither observe the state nor receive any information about it during the game. A central concept for analyzing long-duration stochastic games is the \textit{uniform value}. A game has a uniform value $v$ if for every $\varepsilon>0$, Player 1 (resp., Player 2) has a strategy such that, for all sufficiently large $n$, his average payoff over $n$ stages is at least $v-\varepsilon$ (resp., at most $v+\varepsilon$). Prior work has shown that the uniform value may not exist in general blind stochastic games. To address this, we introduce a subclass called \textit{ergodic blind stochastic games}, defined by imposing an ergodicity condition on the state transitions. For this subclass, we prove the existence of the uniform value and provide an algorithm to approximate it, establishing the \textit{decidability} of the approximation problem. Notably, this decidability result is novel even in the single-player setting of Partially Observable Markov Decision Processes (POMDPs). Furthermore, we show that no algorithm can compute the uniform value exactly, emphasizing the tightness of our result. Finally, we establish that the uniform value is independent of the initial belief.

math.OC