SearcharxivSearch

arXiv subjects

Nicolas Vieille

Publications and source records attributed to Nicolas Vieille.

14 recordsLinked to original sources

Strategic Experimentation with Private Payoffs

We study a strategic experimentation game with exponential bandits, in which experiment outcomes are private. The equilibrium amount of experimentation is always higher than in the benchmark case where experiment outcomes are publicly observed. In addition, for pure equilibria, the equilibrium amount of experimentation is at least socially optimal, and possibly higher. We provide a tight bound on the degree of over-experimentation. The analysis rests on a new form of encouragement effect, according to which a player may hide the absence of a success to encourage future experimentation by the other player, which incentivizes current experimentation.

cs.GT

Undiscounted Equilibrium in Positive Recursive Absorbing Games with Non-Rectangular Absorption Structure

An absorbing game is a stochastic game with a single nonabsorbing state. Such a game is called recursive if all players receive a payoff of 0 in the nonabsorbing state, and positive if all payoffs in absorbing states are positive. An action profile is nonabsorbing if, when it is played, the game remains in the nonabsorbing state with probability 1. The set of nonabsorbing action profiles can be partitioned into the connected components of an undirected graph, whose vertices are these profiles, with two vertices joined by an edge whenever the corresponding profiles differ in the action of a single player. A connected component is said to be rectangular if it is the Cartesian product of subsets of the players' action sets. We prove that every positive recursive absorbing game whose nonabsorbing components are all non-rectangular admits an undiscounted equilibrium payoff.

math.OC

Playing against a stationary opponent

This paper investigates properties of Blackwell $ε$-optimal strategies in zero-sum stochastic games when the adversary is restricted to stationary strategies, motivated by applications to robust Markov decision processes. For a class of absorbing games, we show that Markovian Blackwell $ε$-optimal strategies may fail to exist, yet we prove the existence of Blackwell $ε$-optimal strategies that can be implemented by a two-state automaton whose internal transitions are independent of actions. For more general absorbing games, however, there need not exist Blackwell $ε$-optimal strategies that are independent of the adversary's decisions. Our findings point to a contrast between absorbing games and generalized Big Match games, and provide new insights into the properties of optimal policies for robust Markov decision processes.

cs.GT

Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality

Robust Markov Decision Processes (RMDPs) are a widely used framework for sequential decision-making under parameter uncertainty. RMDPs have been extensively studied when the objective is to maximize the discounted return, but little is known for average optimality (optimizing the long-run average of the rewards obtained over time) and Blackwell optimality (remaining discount optimal for all discount factors sufficiently close to ). In this paper, we prove several foundational results for RMDPs beyond the discounted return. We show that average optimal policies can be chosen stationary and deterministic for sa-rectangular RMDPs but, perhaps surprisingly, we show that for s-rectangular RMDPs average optimal policies may not exist, and if they exist, may need to be history-dependent (Markovian). We also study Blackwell optimality for sa-rectangular RMDPs, where we show that $ε$-Blackwell optimal policies always exist, although Blackwell optimal policies may not exist. We also provide a sufficient condition for their existence, which encompasses virtually any examples from the literature. We then discuss the connection between average and Blackwell optimality, and we describe several algorithms to compute the optimal average return. Interestingly, our approach leverages the connections between RMDPs and stochastic games. Overall, our paper emphasizes the superior practical properties of distance-based sa-rectangular models over s-rectangular models for average and Blackwell optimality.

math.OC

Stationary social learning in a changing environment

We consider social learning in a changing world. Society can remain responsive to state changes only if agents regularly act upon fresh information, which limits the value of social learning. When the state is close to persistent, a consensus whereby most agents choose the same action typically emerges. The consensus action is not perfectly correlated with the state though, because the society exhibits inertia following state changes. Phases of inertia may be longer when signals are more precise, even if agents draw large samples of past actions, as actions then become too correlated within samples, thereby reducing informativeness and welfare.

econ.TH

Optimal Dynamic Information Provision

We study a dynamic model of information provision. A state of nature evolves according to a Markov chain. An informed advisor decides how much information to provide to an uninformed decision maker, so as to influence his short-term decisions. We deal with a stylized class of situations, in which the decision maker has a risky action and a safe action, and the payoff to the advisor only depends on the action chosen by the decision maker. The greedy disclosure policy is the policy which, at each round, minimizes the amount of information being disclosed in that round, under the constraint that it maximizes the current payoff of the advisor. We prove that the greedy policy is optimal in many cases -- but not always.

math.PR

Markov games with frequent actions and incomplete information

We study a two-player, zero-sum, stochastic game with incomplete information on one side in which the players are allowed to play more and more frequently. The informed player observes the realization of a Markov chain on which the payoffs depend, while the non-informed player only observes his opponent's actions. We show the existence of a limit value as the time span between two consecutive stages vanishes; this value is characterized through an auxiliary optimization problem and as the solution of an Hamilton-Jacobi equation.

math.OC

Random Stopping Times in Stopping Problems and Stopping Games

Three notions of random stopping times exist in the literature. We introduce two concepts of equivalence of random stopping times, motivated by optimal stopping problems and stopping games respectively. We prove that these two concepts coincide and that the three notions of random stopping times are equivalent.

math.PR

Dynamic Sender-Receiver Games

We consider a dynamic version of sender-receiver games, where the sequence of states follows an irreducible Markov chain observed by the sender. Under mild assumptions, we provide a simple characterization of the limit set of equilibrium payoffs, as players become very patient. Under these assumptions, the limit set depends on the Markov chain only through its invariant measure. The (limit) equilibrium payoffs are the feasible payoffs that satisfy an individual rationality condition for the receiver, and an incentive compatibility condition for the sender.

math.PR

Strategic Information Exchange

We study a class of two-player repeated games with incomplete information and informational externalities. In these games, two states are chosen at the outset, and players get private information on the pair, before engaging in repeated play. The payoff of each player only depends on his `own' state and on his own action. We study to what extent, and how, information can be exchanged in equilibrium. We prove that provided the private information of each player is valuable for the other player, the set of sequential equilibrium payoffs converges to the set of feasible and individually rational payoffs as players become patient.

math.PR

Lowest Unique Bid Auctions

We consider a class of auctions (Lowest Unique Bid Auctions) that have achieved a considerable success on the Internet. Bids are made in cents (of euro) and every bidder can bid as many numbers as she wants. The lowest unique bid wins the auction. Every bid has a fixed cost, and once a participant makes a bid, she gets to know whether her bid was unique and whether it was the lowest unique. Information is updated in real time, but every bidder sees only what's relevant to the bids she made. We show that the observed behavior in these auctions differs considerably from what theory would prescribe if all bidders were fully rational. We show that the seller makes money, which would not be the case with rational bidders, and some bidders win the auctions quite often. We describe a possible strategy for these bidders.

math.PR

Approximating a sequence of observations by a simple process

Given an arbitrary long but finite sequence of observations from a finite set, we construct a simple process that approximates the sequence, in the sense that with high probability the empirical frequency, as well as the empirical one-step transitions along a realization from the approximating process, are close to that of the given sequence. We generalize the result to the case where the one-step transitions are required to be in given polyhedra.

math.ST

Perturbed Markov Chains

We study irreducible time-homogenous Markov chains with finite state space in discrete time. We obtain results on the sensitivity of the stationary distribution and other statistical quantities with respect to perturbations of the transition matrix. We define a new closeness relation between transition matrices, and use graph-theoretic techniques, in contrast with the matrix analysis techniques previously used.

math.PR