SearcharxivSearch

arXiv subjects

Philippe Jehiel

Publications and source records attributed to Philippe Jehiel.

4 recordsLinked to original sources

Coarse Q-learning: Indifference, Indeterminacy, and Instability

We introduce Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, and feedback from sampled alternatives is pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, we derive the mean-field dynamics and characterize the steady states as smooth analogues of Valuation Equilibria. The model yields novel long-run phenomena in the high payoff-sensitivity limit: depending on the environment, CQL may exhibit multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium at all, with valuations and choice probabilities converging instead to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.

econ.TH

Endogenous clustering and analogy-based expectation equilibrium

Normal-form two-player games are categorized by players into K analogy classes so as to minimize the prediction error about the behavior of the opponent. This results in Clustered Analogy-Based Expectation Equilibria in which strategies are analogy-based expectation equilibria given the analogy partitions and analogy partitions minimize the prediction errors given the strategies. We distinguish between environments with self-repelling analogy partitions in which some mixing over partitions is required and environments with self-attractive partitions in which several analogy partitions can arise, thereby suggesting new channels of belief heterogeneity and equilibrium multiplicity. Various economic applications are discussed.

econ.TH

Auction Design with Data-Driven Misspecifications

We consider auction environments in which at the time of the auction bidders observe signals about their ex-post value. We introduce a model of novice bidders who do not know know the joint distribution of signals and instead build a statistical model relating others' bids to their own ex post value from the data sets accessible from past similar auctions. Crucially, we assume that only ex post values and bids are accessible while signals observed by bidders in past auctions remain private. We consider steady-states in such environments, and importantly we allow for correlation in the signal distribution. We first observe that data-driven bidders may behave suboptimally in classical auctions such as the second-price or first-price auctions whenever there are correlations. Allowing for a mix of rational (or experienced) and data-driven (novice) bidders results in inefficiencies in such auctions, and we show the inefficiency extends to all auction-like mechanisms in which bidders are restricted to submit one-dimensional (real-valued) bids.

econ.TH

Learning to Play Games in Extensive Form by Valuation

A valuation for a player in a game in extensive form is an assignment of numeric values to the players moves. The valuation reflects the desirability moves. We assume a myopic player, who chooses a move with the highest valuation. Valuations can also be revised, and hopefully improved, after each play of the game. Here, a very simple valuation revision is considered, in which the moves made in a play are assigned the payoff obtained in the play. We show that by adopting such a learning process a player who has a winning strategy in a win-lose game can almost surely guarantee a win in a repeated game. When a player has more than two payoffs, a more elaborate learning procedure is required. We consider one that associates with each move the average payoff in the rounds in which this move was made. When all players adopt this learning procedure, with some perturbations, then, with probability 1, strategies that are close to subgame perfect equilibrium are played after some time. A single player who adopts this procedure can guarantee only her individually rational payoff.

cs.LG