SearcharxivSearch

arXiv subjects

Marco Scarsini

Publications and source records attributed to Marco Scarsini.

At least 19 recordsLinked to original sources

Learning When to Automate: Queue Control in Human-AI Service Systems

We study a human-AI service system in which tasks arrive sequentially and are processed through a two-stage architecture: an automated chatbot followed, when necessary, by a human agent. We consider $T$ sequentially arriving tasks, each belonging to one of $K$ heterogeneous types. For each task the decision maker chooses how many resources to allocate to the chatbot, whose type-dependent success probabilities are initially unknown. Tasks not resolved by the chatbot enter type-dependent human-service queues, where they are processed by a human agent with unknown service rates. This model captures a central tradeoff in hybrid service systems: relying more on automation reduces human congestion but increases chatbot costs, while insufficient automation may overload the human agent. We propose the UCB-DPP policy, which combines Upper Confidence Bounds with Drift-Plus-Penalty control to learn the unknown parameters of the system while making queue-aware decisions. We prove that UCB-DPP achieves regret $\widetilde{\mathcal{O}}(K\sqrt{T})$ and guarantees mean-rate stability of the human-service queues. Simulations on synthetic instances show that the proposed policy outperforms natural baselines.

cs.LG

Rigidity and default in production networks

This paper studies the transmission of productivity shocks in general equilibrium production networks, when firms in different sectors operate under informational rigidity and rely on external debt. Rigidity breaks the Modigliani-Miller irrelevance of leverage and may generate default following shocks, even in equilibrium. The economy consists of firms, banks, and consumers. Under proportional shock transmission, we prove that a unique Walrasian rigid equilibrium exists and provide explicit expressions for equilibrium quantities, prices, and interest rates. We show that, on the one hand, Hulten's theorem fails under rigidity, even without leverage. On the other hand, we prove that welfare is smaller than in the first best if and only if both leverage and rigidity exist. The latter increase the total cost of debt and have inflationary effects on the levered sectors, which propagate downstream, and shift consumption and labor upstream. The occurrence of default depends solely on real shocks and the network structure, while the magnitude of the losses depends also on the connectedness of the economy and the cost of debt of the connected sectors. We provide conditions for default cascades to occur and study two examples of default propagation.

econ.TH

Betting on Bets: Anytime-Valid Tests for Stochastic Dominance

How can we monitor, in real time, whether one uncertain prospect has any upside over another? To answer this question, we develop a novel family of sequential, anytime-valid tests for stochastic dominance (SD), a classical and popular notion for comparing entire distribution functions. The problem is distinct from that of testing mean dominance, and it is particularly useful when comparing distributions with similar means or with ordinal outcomes. We first derive powerful, nonparametric e-processes that quantify evidence against the null hypothesis that one prospect is stochastically dominated by another. For first-order SD, these e-processes are based on mixtures of growth-rate optimal e-variables, yielding a test of power one that retains validity under continuous monitoring. We then generalize the approach to sequential testing for higher-order SD and other integral stochastic orders. Empirically, we find that the tests are competitive in power with classical, non-anytime-valid SD tests. Our real-world application examines a controversial phenomenon in baseball analytics, known as the "third-time-through-the-order (3TTO) penalty," viewed as a monitoring problem. We close by sketching the complementary problem of testing whether a prospect has a definite upside, formalizing conditions under which we can derive a powerful anytime-valid test.

stat.ME

Dynamic Wholesale Pricing under Censored-Demand Learning

This paper studies dynamic wholesale pricing and ordering in a two-tier supply chain where firms share POS data and learn about demand from censored demand data. When stockouts occur, unmet demand is unobserved, so the retailer's order quantity affects not only current profits but also the informativeness of future demand signals. This creates a strategic interaction between pricing, ordering, and learning: the manufacturer can influence the pace of learning through wholesale prices, whereas the retailer internalizes the effect of inventory decisions on future information. We analyze a finite-horizon dynamic game in which a manufacturer sets a wholesale price, the retailer then chooses an order quantity, demand is realized, and both firms observe sales. For Weibull demand with a conjugate prior, we extend a dimensionality-reduction approach from single-agent inventory learning models to a strategic supply-chain setting and use it to establish the existence of a Markov perfect equilibrium. For exponential demand, we further show that the equilibrium is unique and admits a recursive characterization. Our numerical analysis shows that public learning can create conflicting incentives in the supply chain: In order to induce larger orders and reduce future censoring, the manufacturer chooses a wholesale price that is lower than a myopic benchmark. By contrast, because of its forward-looking ordering incentive, the retailer may prefer slower learning to avoid strengthening the manufacturer's future wholesale-pricing position.

cs.GT

Strategic Queues with Priority Classes

We consider a strategic M/M/1 queueing model under a first-come-first-served regime, where customers are split into two classes and class $A$ has priority over class $B$. Customers can decide whether to join the queue or balk, and, in case they have joined the queue, whether and when to renege. We study the equilibrium strategies and compare the equilibrium outcome and the social optimum in the two cases where the social optimum is or is not constrained by priority.

cs.GT

Information Design and Full Implementation in Nonatomic Games

This paper studies the implementation of Bayes correlated equilibria in symmetric Bayesian games with nonatomic players, using direct information structures and obedient strategies. The main results demonstrate full implementation in a class of games with negative payoff externalities, such as congestion and Cournot games. Specifically, if the game admits a strictly concave potential in every state, then for every Bayes correlated equilibrium outcome with finite support and rational action distributions, there exists a direct information structure that implements this outcome under all equilibria. When the potential is weakly concave, we show that all equilibria implement the same expected total payoff. Additionally, all Bayes correlated equilibria, including those with infinite support or irrational action distributions, are approximately implemented.

econ.TH

Basins of Attraction in Two-Player Random Ordinal Potential Games

We consider the class of two-person ordinal potential games where each player has the same number of actions $K$. Each game in this class admits at least one pure Nash equilibrium and the best-response dynamics converges to one of these pure Nash equilibria; which one depends on the starting point. So, each pure Nash equilibrium has a basin of attraction. We pick uniformly at random one game from this class and we study the joint distribution of the sizes of the basins of attraction. We provide an asymptotic exact value for the expected basin of attraction of each pure Nash equilibrium, when the number of actions $K$ goes to infinity.

cs.GT

A Characterization of Universally Optimal Queueing Regimes

We consider an M/M/s queueing model in which customers strategically decide, based on the service reward and waiting cost, whether to join upon arrival or balk and, at any time, whether to remain in the queue or renege. Rational strategic behavior yields an equilibrium whose outcome may be socially efficient or inefficient, depending on the queueing regime. Some regimes yield an efficient equilibrium only under precise calibration to the model parameters. Others are universally optimal, meaning that their equilibrium outcome is efficient for all parameter values. Universal optimality is therefore an appealing property for a planner choosing a queueing regime. We characterize the class of universally optimal queueing regimes. A by-product of our characterization is that preemption plays an unavoidable role in universally optimal regimes.

econ.TH

Monotonicity of Equilibria in Nonatomic Congestion Games

This paper studies the monotonicity of equilibrium costs and equilibrium loads in nonatomic congestion games, in response to variations of the demands. The main goal is to identify conditions under which a paradoxical non-monotone behavior can be excluded. In contrast to routing games with a single commodity, where the network topology is the sole determinant factor for monotonicity, for general congestion games with multiple commodities the structure of the strategy sets plays a crucial role. We frame our study in the general setting of congestion games, with a special focus on singleton congestion games, for which we establish the monotonicity of equilibrium loads with respect to every demand. We then provide conditions for comonotonicity of the equilibrium loads, i.e., we investigate when they jointly increase or decrease after variations of the demands. We finally extend our study from singleton congestion games to the larger class of constrained series-parallel congestion games, whose structure is reminiscent of the concept of a series-parallel network.

cs.GT

Phase Transitions of the Price-of-Anarchy Function in Multi-Commodity Routing Games

We consider the behavior of the price of anarchy and equilibrium flows in nonatomic multi-commodity routing games as a function of the traffic demand. We analyze their smoothness with a special attention to specific values of the demand at which the support of the Wardrop equilibrium exhibits a phase transition with an abrupt change in the set of optimal routes. Typically, when such a phase transition occurs, the price of anarchy function has a breakpoint, \ie is not differentiable. We prove that, if the demand varies proportionally across all commodities, then, at a breakpoint, the largest left or right derivatives of the price of anarchy and of the social cost at equilibrium, are associated with the smaller equilibrium support. This proves -- under the assumption of proportional demand -- a conjecture of O'Hare et al. (2016), who observed this behavior in simulations. We also provide counterexamples showing that this monotonicity of the one-sided derivatives may fail when the demand does not vary proportionally, even if it moves along a straight line not passing through the origin.

cs.GT

Strategic Behavior and No-Regret Learning in Queueing Systems

This paper studies a dynamic discrete-time queuing model where at every period players get a new job and must send all their jobs to a queue that has a limited capacity. Players have an incentive to send their jobs as late as possible; however if a job does not exit the queue by a fixed deadline, the owner of the job incurs a penalty and this job is sent back to the player and joins the queue at the next period. Therefore, stability, i.e. the boundedness of the number of jobs in the system, is not guaranteed. We show that if players are myopically strategic, then the system is stable when the penalty is high enough. Moreover, if players use a learning algorithm derived from a typical no-regret algorithm (exponential weight), then the system is stable when penalties are greater than a bound that depends on the total number of jobs in the system.

cs.GT

Best-Response dynamics in two-person random games with correlated payoffs

We consider finite two-player normal form games with random payoffs. Player A's payoffs are i.i.d. from a uniform distribution. Given p in [0, 1], for any action profile, player B's payoff coincides with player A's payoff with probability p and is i.i.d. from the same uniform distribution with probability 1-p. This model interpolates the model of i.i.d. random payoff used in most of the literature and the model of random potential games. First we study the number of pure Nash equilibria in the above class of games. Then we show that, for any positive p, asymptotically in the number of available actions, best response dynamics reaches a pure Nash equilibrium with high probability.

cs.GT

Online Learning in Supply-Chain Games

We study a repeated game between a supplier and a retailer who want to maximize their respective profits without full knowledge of the problem parameters. After characterizing the uniqueness of the Stackelberg equilibrium of the stage game with complete information, we show that even with partial knowledge of the joint distribution of demand and production costs, natural learning dynamics guarantee convergence of the joint strategy profile of supplier and retailer to the Stackelberg equilibrium of the stage game. We also prove finite-time bounds on the supplier's regret and asymptotic bounds on the retailer's regret, where the specific rates depend on the type of knowledge preliminarily available to the players. In the special case when the supplier is not strategic (vertical integration), we prove optimal finite-time regret bounds on the retailer's regret (or, equivalently, the social welfare) when costs and demand are adversarially generated and the demand is censored.

cs.GT

Making the most of your day: online learning for optimal allocation of time

We study online learning for optimal allocation when the resource to be allocated is time. %Examples of possible applications include job scheduling for a computing server, a driver filling a day with rides, a landlord renting an estate, etc. An agent receives task proposals sequentially according to a Poisson process and can either accept or reject a proposed task. If she accepts the proposal, she is busy for the duration of the task and obtains a reward that depends on the task duration. If she rejects it, she remains on hold until a new task proposal arrives. We study the regret incurred by the agent, first when she knows her reward function but does not know the distribution of the task duration, and then when she does not know her reward function, either. This natural setting bears similarities with contextual (one-armed) bandits, but with the crucial difference that the normalized reward associated to a context depends on the whole distribution of contexts.

stat.ML

Social Learning in Nonatomic Routing Games

We consider a discrete-time nonatomic routing game with variable demand and uncertain costs. Given a routing network with single origin and destination, the cost function of each edge depends on some uncertain persistent state parameter. At every period, a random traffic demand is routed through the network according to a Wardrop equilibrium. The realized costs are publicly observed and the public Bayesian belief about the state parameter is updated. We say that there is strong learning when beliefs converge to the truth and weak learning when the equilibrium flow converges to the complete-information flow. We characterize the networks for which learning occurs. We prove that these networks have a series-parallel structure and provide a counterexample to show that learning may fail in non-series-parallel networks.

econ.TH

Correlated Equilibria in Large Anonymous Bayesian Games

We consider multi-population Bayesian games with a large number of players. Each player aims at minimizing a cost function that depends on this player's own action, the distribution of players' actions in all populations, and an unknown state parameter. We study the nonatomic limit versions of these games and introduce the concept of Bayes correlated Wardrop equilibrium, which extends the concept of Bayes correlated equilibrium to nonatomic games. We prove that Bayes correlated Wardrop equilibria are limits of action flows induced by Bayes correlated equilibria of the game with a large finite set of small players. For nonatomic games with complete information admitting a convex potential, we prove that the set of correlated and of coarse correlated Wardrop equilibria coincide with the set of probability distributions over Wardrop equilibria, and that all equilibrium outcomes have the same costs. We get the following consequences. First, all flow distributions of (coarse) correlated equilibria in convex potential games with finitely many players converge to Wardrop equilibria when the weight of each player tends to zero. Second, for any sequence of flows satisfying a no-regret property, its empirical distribution converges to the set of distributions over Wardrop equilibria and the average cost converges to the unique Wardrop cost.

cs.GT

Efficiency of equilibria in games with random payoffs

We consider normal-form games with $n$ players and two strategies for each player, where the payoffs are i.i.d. random variables with some distribution $F$ and we consider issues related to the pure equilibria in the game as the number of players diverges. It is well-known that, if the distribution $F$ has no atoms, the random number of pure equilibria is asymptotically Poisson$(1)$. In the presence of atoms, it diverges. For each strategy profile, we consider the (random) average payoff of the players, called Average Social Utility (ASU). In particular, we examine the asymptotic behavior of the optimum ASU and the one associated to the best and worst pure Nash equilibria and we show that, although these quantities are random, they converge, as $n\to\infty$ to some deterministic quantities.

math.PR

The Buck-Passing Game

We consider a game in which players are the vertices of a directed graph. Initially, Nature chooses one player according to some fixed distribution and gives her a buck, which represents the request to perform a chore. After completing the task, the player passes the buck to one of her out-neighbors in the graph. The procedure is repeated indefinitely and each player's cost is the asymptotic expected frequency of times that she receives the buck. We consider a deterministic and a stochastic version of the game depending on how players select the neighbor to pass the buck. In both cases we prove the existence of pure equilibria that do not depend on the initial distribution; this is achieved by showing the existence of a generalized ordinal potential. We then use the price of anarchy and price of stability to measure fairness of these equilibria. We also study a buck-holding variant of the game in which players want to maximize the frequency of times they hold the buck, which includes the PageRank game as a special case.

cs.GT