SearcharxivSearch

arXiv subjects

Xavier Venel

Publications and source records attributed to Xavier Venel.

At least 19 recordsLinked to original sources

Dynamic Wholesale Pricing under Censored-Demand Learning

This paper studies dynamic wholesale pricing and ordering in a two-tier supply chain where firms share POS data and learn about demand from censored demand data. When stockouts occur, unmet demand is unobserved, so the retailer's order quantity affects not only current profits but also the informativeness of future demand signals. This creates a strategic interaction between pricing, ordering, and learning: the manufacturer can influence the pace of learning through wholesale prices, whereas the retailer internalizes the effect of inventory decisions on future information. We analyze a finite-horizon dynamic game in which a manufacturer sets a wholesale price, the retailer then chooses an order quantity, demand is realized, and both firms observe sales. For Weibull demand with a conjugate prior, we extend a dimensionality-reduction approach from single-agent inventory learning models to a strategic supply-chain setting and use it to establish the existence of a Markov perfect equilibrium. For exponential demand, we further show that the equilibrium is unique and admits a recursive characterization. Our numerical analysis shows that public learning can create conflicting incentives in the supply chain: In order to induce larger orders and reduce future censoring, the manufacturer chooses a wholesale price that is lower than a myopic benchmark. By contrast, because of its forward-looking ordering incentive, the retailer may prefer slower learning to avoid strengthening the manufacturer's future wholesale-pricing position.

cs.GT

Optimal strategies in Markov decision processes with finitely additive evaluations

We study infinite-horizon Markov decision processes (MDPs) where the decision maker evaluates each of her strategies by aggregating the infinite stream of expected stage-rewards. The crucial feature of our approach is that the aggregation is performed by means of a given diffuse charge (a diffuse finitely additive probability measure) on the set of stages. The results of Neyman [2023] imply that in this setting, in every MDP with finite state and action spaces, the decision maker has a pure optimal strategy as long as the diffuse charge satisfies the time value of money principle. His result raises the question of existence of an optimal strategy without additional assumptions on the aggregation charge. We answer this question in the negative with a counterexample. With a delicately constructed aggregation charge, the MDP has no optimal strategy at all, neither pure nor randomized.

math.OC

Continuous Social Networks

We develop an extension of the classical model of DeGroot (1974) to a continuum of agents when they interact among them according to a DiKernel. We show that, under some regularity assumptions, the continuous model is the limit case of the discrete one. Additionally, we establish sufficient conditions for the emergence of consensus. We provide some applications of these results. First, we establish a canonical way to reduce the dimensionality of matrices by comparing matrices of different dimensions in the space of DiKernels. Then, we develop a model of Lobby Competition where two lobbies compete to bias the opinion of a continuum of agents. We give sufficient conditions for the existence of a Nash Equilibrium and study their relation with the equilibria of discretizations of the game. Finally, we characterize the equilibrium for a particular case of DiKernels.

econ.TH

Comparing experiments in discounted problems

This paper compares statistical experiments in discounted problems, ranging from the simplest ones where the state is fixed and the flow of information exogenous to more complex ones, where the decision-maker controls the flow of information or the state changes over time.

econ.TH

Zero-one Laws for a Control Problem with Random Action Sets

In many control problems there is only limited information about the actions that will be available at future stages. We introduce a framework where the Controller chooses actions $a_{0}, a_{1}, \ldots$, one at a time. Her goal is to maximize the probability that the infinite sequence $(a_{0}, a_{1}, \ldots)$ is an element of a given subset $G$ of $\mathbb{N}^{\mathbb{N}}$. The set $G$, called the goal, is assumed to be a Borel tail set. The Controller's choices are restricted: having taken a sequence $h_{t} = (a_{0}, \ldots, a_{t-1})$ of actions prior to stage $t \in \mathbb{N}$, she must choose an action $a_{t}$ at stage $t$ from a non-empty, finite subset $A(h_{t})$ of $\mathbb{N}$. The set $A(h_{t})$ is chosen from a distribution $p_{t}$, independently over all $t \in \mathbb{N}$ and all $h_{t} \in \mathbb{N}^{t}$. We consider several information structures defined by how far ahead into the future the Controller knows what actions will be available. In the special case where all the action sets are singletons (and thus the Controller is a dummy), Kolmogorov's 0-1 law says that the probability for the goal to be reached is 0 or 1. We construct a number of counterexamples to show that in general the value of the control problem can be strictly between 0 and 1, and derive several sufficient conditions for the 0-1 ``law" to hold.

math.OC

Weighted Average-convexity and Cooperative Games

We generalize the notion of convexity and average-convexity to the notion of weighted average-convexity. We show several results on the relation between weighted average-convexity and cooperative games. First, we prove that if a game is weighted average-convex, then the corresponding weighted Shapley value is in the core. Second, we exhibit necessary conditions for a communication TU-game to preserve the weighted average-convexity. Finally, we provide a complete characterization when the underlying graph is a priority decreasing tree.

cs.GT

Strategic Behavior and No-Regret Learning in Queueing Systems

This paper studies a dynamic discrete-time queuing model where at every period players get a new job and must send all their jobs to a queue that has a limited capacity. Players have an incentive to send their jobs as late as possible; however if a job does not exit the queue by a fixed deadline, the owner of the job incurs a penalty and this job is sent back to the player and joins the queue at the next period. Therefore, stability, i.e. the boundedness of the number of jobs in the system, is not guaranteed. We show that if players are myopically strategic, then the system is stable when the penalty is high enough. Moreover, if players use a learning algorithm derived from a typical no-regret algorithm (exponential weight), then the system is stable when penalties are greater than a bound that depends on the total number of jobs in the system.

cs.GT

Repeated Games with Switching Costs: Stationary vs History Independent Strategies

We study zero-sum repeated games where the minimizing player has to pay a certain cost each time he changes his action. Our contribution is twofold. First, we show that the value of the game exists in stationary strategies, depending solely on the previous action of the minimizing player, not the entire history. We provide a full characterization of the value and the optimal strategies. The strategies exhibit a robustness property and typically do not change with a small perturbation of the switching costs. Second, we consider a case where the minimizing player is limited to playing simpler strategies that are completely history-independent. Here too, we provide a full characterization of the (minimax) value and the strategies for obtaining it. Moreover, we present several bounds on the loss due to this limitation.

math.OC

Diffusion in large networks

We investigate the phenomenon of diffusion in a countably infinite society of individuals interacting with their neighbors in a network. At a given time, each individual is either active or inactive. The diffusion is driven by two characteristics: the network structure and the diffusion mechanism represented by an aggregation function. We distinguish between two diffusion mechanisms (probabilistic, deterministic) and focus on two types of aggregation functions (strict, Boolean). Under strict aggregation functions, polarization of the society cannot happen, and its state evolves towards a mixture of infinitely many active and infinitely many inactive agents, or towards a homogeneous society. Under Boolean aggregation functions, the diffusion process becomes deterministic and the contagion model of Morris (2000) becomes a particular case of our framework. Polarization can then happen. Our dynamics also allows for cycles in both cases. The network structure is not relevant for these questions, but is important for establishing irreducibility, at the price of a richness assumption: the network should contain infinitely many complex stars and have enough space for storing local configurations. Our model can be given a game-theoretic interpretation via a local coordination game, where each player would apply a best-response strategy in a random neighborhood.

cs.SI

Robust communication on networks

We consider sender-receiver games, where the sender and the receiver are two distinct nodes in a communication network. Communication between the sender and the receiver is thus indirect. We ask when it is possible to robustly implement the equilibrium outcomes of the direct communication game as equilibrium outcomes of indirect communication games on the network. Robust implementation requires that: (i) the implementation is independent of the preferences of the intermediaries and (ii) the implementation is guaranteed at all histories consistent with unilateral deviations by the intermediaries. Robust implementation of direct communication is possible if and only if either the sender and receiver are directly connected or there exist two disjoint paths between the sender and the receiver.

econ.TH

History-dependent evaluations in POMDPs

We consider POMDPs in which the weight of the stage payoff depends on the past sequence of signals and actions occurring in the infinitely repeated problem. We prove that for all epsilon>0, there exists a strategy that is epsilon-optimal for any sequence of weights satisfying a property that interprets as "the decision-maker is patient enough". This unifies and generalizes several results of the literature, and applies notably to POMDPs with limsup payoffs.

math.OC

Decomposition of games: some strategic considerations

Candogan et al. (2011) provide an orthogonal direct-sum decomposition of finite games into potential, harmonic and nonstrategic components. In this paper we study the issue of decomposing games that are strategically equivalent from a game-theoretical point of view, for instance games obtained via transformations such as duplications of strategies or positive affine mappings of of payoffs. We show the need to define classes of decompositions to achieve commutativity of game transformations and decompositions.

cs.GT

Commutative Stochastic Games

We are interested in the convergence of the value of n-stage games as n goes to infinity and the existence of the uniform value in stochastic games with a general set of states and finite sets of actions where the transition is commutative. This means that playing an action profile a 1 followed by an action profile a 2 , leads to the same distribution on states as playing first the action profile a 2 and then a 1. For example, absorbing games can be reformulated as commutative stochastic games. When there is only one player and the transition function is deterministic, we show that the existence of a uniform value in pure strategies implies the existence of 0-optimal strategies. In the framework of two-player stochastic games, we study a class of games where the set of states is R m and the transition is deterministic and 1-Lipschitz for the L 1-norm, and prove that these games have a uniform value. A similar proof shows the existence of an equilibrium in the non zero-sum case. These results remain true if one considers a general model of finite repeated games, where the transition is commutative and the players observe the past actions but not the state.

math.OC

On values of repeated games with signals

We study the existence of different notions of value in two-person zero-sum repeated games where the state evolves and players receive signals. We provide some examples showing that the limsup value (and the uniform value) may not exist in general. Then we show the existence of the value for any Borel payoff function if the players observe a public signal including the actions played. We also prove two other positive results without assumptions on the signaling structure: the existence of the $\sup$ value in any game and the existence of the uniform value in recursive games with nonnegative payoffs.

math.OC

Pathwise uniform value in gambling houses and Partially Observable Markov Decision Processes

In several standard models of dynamic programming (gambling houses, MDPs, POMDPs), we prove the existence of a very robust notion of value for the infinitely repeated problem, namely the pathwise uniform value. This solves two open problems. First, this shows that for any epsilon>0, the decision-maker has a pure strategy sigma which is epsilon-optimal in any n-stage game, provided that n is big enough (this result was only known for behavior strategies, that is, strategies which use randomization). Second, the strategy sigma can be chosen such that under the long-run average payoff criterion (expectation of the liminf of the average payoffs), the decision-maker has more than lim v(n)-epsilon.

math.OC

Recursive games: Uniform value, Tauberian theorem and the Mertens conjecture "$Maxmin=\lim v_n=\lim v_λ$"

We study two-player zero-sum recursive games with a countable state space and finite action spaces at each state. When the family of $n$-stage values $\{v_n,n\geq 1\}$ is totally bounded for the uniform norm, we prove the existence of the uniform value. Together with a result in Rosenberg and Vieille (2000), we obtain a uniform Tauberian theorem for recursive games: $(v_n)$ converges uniformly if and only if $(v_λ)$ converges uniformly. We apply our main result to finite recursive games with signals (where players observe only signals on the state and on past actions). When the maximizer is more informed than the minimizer, we prove the Mertens conjecture $Maxmin=\lim_{n\to\infty} v_n=\lim_{λ\to 0}v_λ$. Finally, we deduce the existence of the uniform value in finite recursive game with symmetric information.

math.OC

Attainability in Repeated Games with Vector Payoffs

We introduce the concept of attainable sets of payoffs in two-player repeated games with vector payoffs. A set of payoff vectors is called {\em attainable} if player 1 can ensure that there is a finite horizon $T$ such that after time $T$ the distance between the set and the cumulative payoff is arbitrarily small, regardless of what strategy player 2 is using. This paper focuses on the case where the attainable set consists of one payoff vector. In this case the vector is called an attainable vector. We study properties of the set of attainable vectors, and characterize when a specific vector is attainable and when every vector is attainable.

math.OC

Existence of the uniform value in repeated games with a more informed controller

We prove that in a general zero-sum repeated game where the first player is more informed than the second player and controls the evolution of information on the state, the uniform value exists. This result extends previous results on Markov decision processes with partial observation (Rosenberg, Solan, Vieille 2002), and repeated games with an informed controller (Renault 2012). Our formal definition of a more informed player is more general than the inclusion of signals, allowing therefore for imperfect monitoring of actions. We construct an auxiliary stochastic game whose state space is the set of second order beliefs of player 2 (beliefs about beliefs of player 1 on the true state variable of the initial game) with perfect monitoring and we prove it has a value by using a result of Renault 2012. A key element in this work is to prove that player 1 can use strategies of the auxiliary game in the initial game in our general framework, which allows to deduce that the value of the auxiliary game is also the value of our initial repeated game by using classical arguments.

math.OC