SearcharxivSearch

arXiv subjects

Bruno Ziliotto

Publications and source records attributed to Bruno Ziliotto.

At least 19 recordsLinked to original sources

Approximating the Uniform Value in Hidden Stochastic Games with Doeblin Condition

We study \emph{zero-sum two-player hidden stochastic games}, where players receive partial observations of the state. We focus on a central solution concept for analyzing long-duration stochastic games: the \emph{uniform value}, a limiting average payoff that both players can guarantee for sufficiently long durations. In the general case, prior work provides examples of games that do not have a uniform value. Moreover, for the subclass of games that do have a uniform value, there exists no algorithm that approximates it. Therefore, we generalize the \emph{Doeblin condition} for Markov chains (which guarantees the existence of a unique invariant measure) to hidden stochastic games. Informally, the Doeblin condition for hidden stochastic games requires that, for every way to play the game, there exists a fixed belief such that, no matter the initial belief over the state of the game, after sufficiently many stages, the posterior belief is probably close to this fixed belief. Under the Doeblin condition, we prove the existence of the uniform value, provide an algorithm to approximate it, and prove that no algorithm can compute it exactly. Then, we identify structural conditions on the transition function that ensure the Doeblin condition holds both in the blind setting, where observations are uninformative, and in the hidden setting, where observations are partially informative. When considering games with only one player, namely partially observable Markov decision processes, our results provide a novel subclass in which the uniform value exists and can be approximated, but cannot be computed exactly

math.OC

The Role of Commitment in Optimal Stopping

We investigate the role of commitment in optimal stopping by studying all the variants between Prophet Inequality (PI) and Pandora's Box (PB). Both problems deal with a set of variables drawn from known distributions. In PI the gambler observes an adversarial order of these variables with the goal of selecting one that maximizes the expected value against a prophet who knows the exact values realized. The gambler has to irrevocably decide at each step whether to select the value or discard it (commitment). On the other hand, in PB the gambler selects the order of inspecting the variables and for each pays an observation cost to see the actual value realized, aiming to choose one to maximize the net cost of the value chosen minus the observation cost paid. The gambler in PB can return and select any variable already seen (no commitment). For all the variants between these problems that arise by changing parameters such as (1) commitment (2) observation cost (3) order selection, we concisely summarize the known results and fill the gaps of variants not yet studied. We also uncover connections to Ski-Rental, a classic online algorithm problem.

cs.DS

Residual Prophet Inequalities

We introduce a variant of the classic prophet inequality, called \emph{residual prophet inequality} (RPI). In the RPI problem, we consider a finite sequence of $n$ nonnegative independent random values with known distributions, and a known integer $0\leq k\leq n-1$. Before the gambler observes the sequence, the top $k$ values are removed, whereas the remaining $n-k$ values are streamed sequentially to the gambler. For example, one can assume that the top $k$ values have already been allocated to a higher-priority agent. Upon observing a value, the gambler must decide irrevocably whether to accept or reject it, without the possibility of revisiting past values. We study two variants of RPI, according to whether the gambler learns online of the identity of the variable that he sees (FI model) or not (NI model). Our main result is a randomized algorithm in the FI model with \emph{competitive ratio} of at least $1/(k+2)$, which we show is tight. Our algorithm is data-driven and requires access only to the $k+1$ largest values of a single sample from the $n$ input distributions. In the NI model, we provide a similar algorithm that guarantees a competitive ratio of $1/(2k+2)$. We further analyze independent and identically distributed instances when $k=1$. We build a single-threshold algorithm with a competitive ratio of at least 0.4901, and show that no single-threshold strategy can get a competitive ratio greater than 0.5464.

cs.DS

Constant Payoff Property in Zero-Sum Stochastic Games with a Finite Horizon

This paper examines finite zero-sum stochastic games and demonstrates that when the game's duration is sufficiently long, there exists a pair of approximately optimal strategies such that the expected average payoff at any point in the game remains close to the value. This property, known as the \textit{constant payoff property}, was previously established only for absorbing games and discounted stochastic games.

math.OC

The game behind oriented percolation

We characterize the critical parameter of oriented percolation on $\mathbb{Z}^2$ through the value of a zero-sum game. Specifically, we define a zero-sum game on a percolation configuration of $\mathbb{Z}^2$, where two players move a token along the non-oriented edges of $\mathbb{Z}^2$, collecting a cost of 1 for each edge that is open, and 0 otherwise. The total cost is given by the limit superior of the average cost. We demonstrate that the value of this game is deterministic and equals 1 if and only if the percolation parameter exceeds $p_c$, the critical exponent of oriented percolation. Additionally, we establish that the value of the game is continuous at $p_c$. Finally, we show that for $p$ close to 0, the value of the game is equal to 0.

math.PR

Time-Dependent Blackwell Approachability and Application to Absorbing Games

Blackwell's approachability (Blackwell, 1954, 1956) is a very general online learning framework where a Decision Maker obtains vector-valued outcomes, and aims at the convergence of the average outcome to a given ``target'' set. Blackwell gave a sufficient condition for the decision maker having a strategy guaranteeing such a convergence against an adversarial environment, as well as what we now call the Blackwell's algorithm, which then ensures convergence. Blackwell's approachability has since been applied to numerous problems, in regret minimization and game theory, in particular. We extend this framework by allowing the outcome function and the inner product to be time-dependent. We establish a general guarantee for the natural extension to this framework of Blackwell's algorithm. In the case where the target set is an orthant, we present a family of time-dependent inner products which yields different convergence speeds for each coordinate of the average outcome. We apply this framework to absorbing games (an important class of stochastic games) for which we construct $\varepsilon$-uniformly optimal strategies using Blackwell's algorithm in a well-chosen auxiliary approachability problem, thereby giving a novel illustration of the relevance of online learning tools for solving games.

math.OC

Stochastic Homogenization of HJ Equations: a Differential Game Approach

We prove stochastic homogenization for a class of non-convex and non-coercive first-order Hamilton-Jacobi equations in a finite-range-dependence environment for Hamiltonians that can be expressed by a max-min formula. Exploiting the representation of solutions as value functions of differential games, we develop a game-theoretic approach to homogenization. We furthermore extend this result to a class of Lipschitz Hamiltonians that need not admit a global max-min representation. Our methods allow us to get a quantitative convergence rate for solutions with linear initial data toward the corresponding ones of the effective limit problem.

math.AP

Uniform Value and Decidability in Ergodic Blind Stochastic Games

We study a class of two-player zero-sum stochastic games known as \textit{blind stochastic games}, where players neither observe the state nor receive any information about it during the game. A central concept for analyzing long-duration stochastic games is the \textit{uniform value}. A game has a uniform value $v$ if for every $\varepsilon>0$, Player 1 (resp., Player 2) has a strategy such that, for all sufficiently large $n$, his average payoff over $n$ stages is at least $v-\varepsilon$ (resp., at most $v+\varepsilon$). Prior work has shown that the uniform value may not exist in general blind stochastic games. To address this, we introduce a subclass called \textit{ergodic blind stochastic games}, defined by imposing an ergodicity condition on the state transitions. For this subclass, we prove the existence of the uniform value and provide an algorithm to approximate it, establishing the \textit{decidability} of the approximation problem. Notably, this decidability result is novel even in the single-player setting of Partially Observable Markov Decision Processes (POMDPs). Furthermore, we show that no algorithm can compute the uniform value exactly, emphasizing the tightness of our result. Finally, we establish that the uniform value is independent of the initial belief.

math.OC

Bayesian Learning in Mean Field Games

We consider a mean-field game model where the cost functions depend on a fixed parameter, called \textit{state}, which is unknown to players. Players learn about the state from a a stream of private signals they receive throughout the game. We derive a mean field system satisfied by the equilibrium payoff of the game and prove existence of a solution under standard regularity assumptions. Additionally, we establish the uniqueness of the solution when the cost function satisfies the monotonicity assumption of Lasry and Lions at each state.

math.OC

Zero-sum Random Games on Directed Graphs

This paper considers a class of two-player zero-sum games on directed graphs whose vertices are equipped with random payoffs of bounded support known by both players. Starting from a fixed vertex, players take turns to move a token along the edges of the graph. On the one hand, for acyclic directed graphs of bounded degree and sub-exponential expansion, we show that the value of the game converges almost surely to a constant at an exponential rate dominated in terms of the expansion. On the other hand, for the infinite $d$-ary tree that does not fall into the previous class of graphs, we show convergence at a double-exponential rate in terms of the expansion.

math.OC

Prophet Inequalities Require Only a Constant Number of Samples

In a prophet inequality problem, $n$ independent random variables are presented to a gambler one by one. The gambler decides when to stop the sequence and obtains the most recent value as reward. We evaluate a stopping rule by the worst-case ratio between its expected reward and the expectation of the maximum variable. In the classic setting, the order is fixed, and the optimal ratio is known to be 1/2. Three variants of this problem have been extensively studied: the prophet-secretary model, where variables arrive in uniformly random order; the free-order model, where the gambler chooses the arrival order; and the i.i.d. model, where the distributions are all the same, rendering the arrival order irrelevant. Most of the literature assumes that distributions are known to the gambler. Recent work has considered the question of what is achievable when the gambler has access only to a few samples per distribution. Surprisingly, in the fixed-order case, a single sample from each distribution is enough to approximate the optimal ratio, but this is not the case in any of the three variants. We provide a unified proof that for all three variants of the problem, a constant number of samples (independent of n) for each distribution is good enough to approximate the optimal ratios. Prior to our work, this was known to be the case only in the i.i.d. variant. We complement our result showing that our algorithms can be implemented in polynomial time. A key ingredient in our proof is an existential result based on a minimax argument, which states that there must exist an algorithm that attains the optimal ratio and does not rely on the knowledge of the upper tail of the distributions. A second key ingredient is a refined sample-based version of a decomposition of the instance into "small" and "large" variables, first introduced by Liu et al. [EC'21].

cs.DS

Finite-Memory Strategies in POMDPs with Long-Run Average Objectives

Partially observable Markov decision processes (POMDPs) are standard models for dynamic systems with probabilistic and nondeterministic behaviour in uncertain environments. We prove that in POMDPs with long-run average objective, the decision maker has approximately optimal strategies with finite memory. This implies notably that approximating the long-run value is recursively enumerable, as well as a weak continuity property of the value with respect to the transition function.

cs.GT

Constant payoff in zero-sum stochastic games

In a zero-sum stochastic game, at each stage, two adversary players take decisions and receive a stage payoff determined by them and by a controlled random variable representing the state of nature. The total payoff is the normalized discounted sum of the stage payoffs. In this paper we solve the "constant payoff" conjecture formulated by Sorin, Vigeral and Venel (2010): if both players use optimal strategies, then for any alpha>0, the expected discounted payoff between stage 1 and stage alpha/lambda tends to the limit discounted value of the game, as the discount rate lambda goes to 0.

math.OC

Percolation games

This paper introduces a discrete-time stochastic game class on $\mathbb{Z}^d$, which plays the role of a toy model for the well-known problem of stochastic homogenization of Hamilton-Jacobi equations. Conditions are provided under which the $n$-stage game value converges as $n$ tends to infinity, and connections with homogenization theory is discussed.

math.OC

Mertens conjectures in absorbing games with incomplete information

In a zero-sum stochastic game with signals, at each stage, two adversary players take decisions and receive a stage payoff determined by these decisions and a variable called state. The state follows a Markov chain, that is controlled by both players. Actions and states are imperfectly observed by players, who receive a private signal at each stage. Mertens (ICM 1986) conjectured two properties regarding games with long duration: first, that limit value always exists, second, that when Player 1 is more informed than Player 2, she can guarantee uniformly the limit value. These conjectures were disproved recently by the author, but remain widely open in many subclasses. A well-known particular subclass is the one of absorbing games with incomplete information on both sides, in which the state can move at most once during the game, and players get a private signal about it at the outset of the game. This paper proves Mertens conjectures in this particular model, by introducing a new approximation technique of belief dynamics, that is likely to generalize to many other frameworks. In particular, this makes a significant step towards the understanding of the following broad question: in which games do Mertens conjectures hold?

math.OC

Unknown I.I.D. Prophets: Better Bounds, Streaming Algorithms, and a New Impossibility

A prophet inequality states, for some $α\in[0,1]$, that the expected value achievable by a gambler who sequentially observes random variables $X_1,\dots,X_n$ and selects one of them is at least an $α$ fraction of the maximum value in the sequence. We obtain three distinct improvements for a setting that was first studied by Correa et al. (EC, 2019) and is particularly relevant to modern applications in algorithmic pricing. In this setting, the random variables are i.i.d. from an unknown distribution and the gambler has access to an additional $βn$ samples for some $β\geq 0$. We first give improved lower bounds on $α$ for a wide range of values of $β$; specifically, $α\geq(1+β)/e$ when $β\leq 1/(e-1)$, which is tight, and $α\geq 0.648$ when $β=1$, which improves on a bound of around $0.635$ due to Correa et al. (SODA, 2020). Adding to their practical appeal, specifically in the context of algorithmic pricing, we then show that the new bounds can be obtained even in a streaming model of computation and thus in situations where the use of relevant data is complicated by the sheer amount of data available. We finally establish that the upper bound of $1/e$ for the case without samples is robust to additional information about the distribution, and applies also to sequences of i.i.d. random variables whose distribution is itself drawn, according to a known distribution, from a finite set of known candidate distributions. This implies a tight prophet inequality for exchangeable sequences of random variables, answering a question of Hill and Kertz (Contemporary Mathematics, 1992), but leaves open the possibility of better guarantees when the number of candidate distributions is small, a setting we believe is of strong interest to applications.

cs.DS

History-dependent evaluations in POMDPs

We consider POMDPs in which the weight of the stage payoff depends on the past sequence of signals and actions occurring in the infinitely repeated problem. We prove that for all epsilon>0, there exists a strategy that is epsilon-optimal for any sequence of weights satisfying a property that interprets as "the decision-maker is patient enough". This unifies and generalizes several results of the literature, and applies notably to POMDPs with limsup payoffs.

math.OC