SearcharxivSearch

arXiv subjects

Yifen Mu

Publications and source records attributed to Yifen Mu.

10 recordsLinked to original sources

Robust Aggregation of Calibrated Forecasts

Decision-makers often rely on multiple probabilistic forecasts that are individually calibrated but need not be fully informative. We develop a framework for aggregating such forecasts when the decision-maker knows only that experts satisfy calibration. We show that the joint distribution of calibrated forecasts can contain decision-relevant information that is unavailable from any single expert, so the standard optimal-in-hindsight (OIH) benchmark may substantially understate attainable performance. To formalize this idea, we introduce a robust max-min benchmark: the best payoff a decision-maker can guarantee against all profile-wise conditional-mean mappings compatible with calibration. This benchmark is tractable, admits a linear-programming formulation, and dominates the OIH benchmark up to calibration error. It can nevertheless be strictly below the Bayesian benchmark, clarifying the value of knowing experts' information structures. Finally, we provide online algorithms that attain the robust benchmark under forecast-only feedback and stronger contextual benchmarks under state feedback.

econ.TH

Decentralized MARL for Coarse Correlated Equilibrium in Aggregative Markov Games

This paper studies the problem of decentralized learning of Coarse Correlated Equilibrium (CCE) in aggregative Markov games (AMGs), where each agent's instantaneous reward depends only on its own action and an aggregate quantity. Existing CCE learning algorithms for general Markov games are not designed to leverage the aggregative structure, and research on decentralized CCE learning for AMGs remains limited. We propose an adaptive stage-based V-learning algorithm that exploits the aggregative structure under a fully decentralized information setting. Based on the two-timescale idea, the algorithm partitions learning into stages and adjusts stage lengths based on the variability of aggregate signals, while using no-regret updates within each stage. We prove the algorithm achieves an epsilon-approximate CCE in O(S Amax T5 / epsilon2) episodes, avoiding the curse of multiagents which commonly arises in MARL. Numerical results verify the theoretical findings, and the decentralized, model-free design enables easy extension to large-scale multi-agent scenarios.

cs.GT

Private Markovian Equilibrium in Stackelberg Markov Games for Smart Grid Demand Response

The increasing integration of renewable energy introduces a great challenge to the supply and demand balance of the power grid. To address this challenge, this paper formulates a Stackelberg Markov game (SMG) between an aggregator and multiple users, where the aggregator sets electricity prices and users make demand and storage decisions. Considering that users' storage levels are private information, we introduce private states and propose the new concepts of private Markovian strategies (PMS) and private Markovian equilibrium (PME). We establish the existence of a pure PME in the lower-level Markov game and prove that it can be computed in polynomial time. Notably, computing equilibrium in general Markov games is hard, and polynomial-time algorithms are rarely available. Based on these theoretical results, we develop a scalable solution framework combining centralized and decentralized algorithms for the lower-level PME computation with upper-level pricing optimization. Numerical simulations with up to 50 users based on real data validate the effectiveness and scalability of the proposed methods, whereas prior studies typically consider no more than 5 users.

eess.SY

On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ games

In infinite-horizon discrete-time linear-quadratic (LQ) dynamic games, computing feedback Nash equilibria (FNEs) remains computationally challenging. Motivated by this, we study a finite-horizon strategy for approximating one of the infinite-horizon FNEs. The finite-horizon strategy is as follows. Each player $i$ has an individual prediction horizon $T^i$. In the infinite-horizon game, at each stage, each player $i$ computes its control in the following way: player $i$ envisions an auxiliary $T^i$-stage game in which the same set of players play, computes the unique FNE of the auxiliary game using a standard method, and implements only the first-stage control. Our main result is, under suitable conditions, the total cost under these finite-horizon strategies converges to that under one of the infinite-horizon FNEs when all players' prediction horizons tend to infinity. Moreover, we derive an explicit cubic-polynomial upper bound on this cost gap with respect to the distance between the corresponding strategy matrices. This strategy is tractable and implementable, as it avoids the direct solution of the coupled algebraic Riccati equations (CARE) of infinite-horizon LQ games.

eess.SY

On the convergence of fictitious play algorithm in repeated games via the geometrical approach

As the earliest and one of the most fundamental learning dynamics for computing NE, fictitious play (FP) has being receiving incessant research attention and finding games where FP would converge (games with FPP) is one central question in related fields. In this paper, we identify a new class of games with FPP, i.e., $3\times3$ games without IIP, based on the geometrical approach by leveraging the location of NE and the partition of best response region. During the process, we devise a new projection mapping to reduce a high-dimensional dynamical system to a planar system. And to overcome the non-smoothness of the systems, we redefine the concepts of saddle and sink NE, which are proven to exist and help prove the convergence of CFP by separating the projected space into two parts. Furthermore, we show that our projection mapping can be extended to higher-dimensional and degenerate games.

math.OC

On Game based Distributed Approach for General Multi-agent Optimal Coverage with Application to UAV Networks

This paper focuses on the optimal coverage problem (OCP) for multi-agent systems with a decentralized optimization mechanism. A game based distributed decision-making method for the multi-agent OCP is proposed to address the high computational costs arising from the large scale of the multi-agent system and to ensure that the game's equilibrium achieves the global performance objective's maximum value. In particular, a distributed algorithm that needs only local information is developed and proved to converge to near-optimal global coverage. Finally, the proposed method is applied to maximize the coverage area of the UAV network for a target region. The simulation results show that our method can require much less computational time than other typical distributed algorithms in related work, while achieving a faster convergence rate. Comparison with centralized optimization also demonstrates that the proposed method has approximate optimization results and high computation efficiency.

eess.SY

An Optimal Pricing Formula for Smart Grid based on Stackelberg Game

The dynamic pricing of electricity is one of the most crucial demand response (DR) strategies in smart grid, where the utility company typically adjust electricity prices to influence user electricity demand. This paper models the relationship between the utility company and flexible electricity users as a Stackelberg game. Based on this model, we present a series of analytical results under certain conditions. First, we give an analytical Stackelberg equilibrium, namely the optimal pricing formula for utility company, as well as the unique and strict Nash equilibrium for users' electricity demand under this pricing scheme. To our best knowledge, it is the first optimal pricing formula in the research of price-based DR strategies. Also, if there exist prediction errors for the supply and demand of electricity, we provide an analytical expression for the energy supply cost of utility company. Moreover, a sufficient condition has been proposed that all electricity demands can be supplied by renewable energy. When the conditions for analytical results are not met, we provide a numerical solution algorithm for the Stackelberg equilibrium and verify its efficiency by simulation.

math.OC

A Payoff-Based Policy Gradient Method in Stochastic Games with Long-Run Average Payoffs

Despite the significant potential for various applications, stochastic games with long-run average payoffs have received limited scholarly attention, particularly concerning the development of learning algorithms for them due to the challenges of mathematical analysis. In this paper, we study the stochastic games with long-run average payoffs and present an equivalent formulation for individual payoff gradients by defining advantage functions which will be proved to be bounded. This discovery allows us to demonstrate that the individual payoff gradient function is Lipschitz continuous with respect to the policy profile and that the value function of the games exhibits the gradient dominance property. Leveraging these insights, we devise a payoff-based gradient estimation approach and integrate it with the Regularized Robbins-Monro method from stochastic approximation theory to construct a bandit learning algorithm suited for stochastic games with long-run average payoffs. Additionally, we prove that if all players adopt our algorithm, the policy profile employed will asymptotically converge to a Nash equilibrium with probability one, provided that all Nash equilibria are globally neutrally stable and a globally variationally stable Nash equilibrium exists. This condition represents a wide class of games, including monotone games.

cs.GT

Periodicity in Hedge-myopic system and an asymmetric NE-solving paradigm for two-player zero-sum games

In this paper, we consider the $n \times n$ two-payer zero-sum repeated game in which one player (player X) employs the popular Hedge (also called multiplicative weights update) learning algorithm while the other player (player Y) adopts the myopic best response. We investigate the dynamics of such Hedge-myopic system by defining a metric $Q(\textbf{x}_t)$, which measures the distance between the stage strategy $\textbf{x}_t$ and Nash Equilibrium (NE) strategy of player X. We analyze the trend of $Q(\textbf{x}_t)$ and prove that it is bounded and can only take finite values on the evolutionary path when the payoff matrix is rational and the game has an interior NE. Based on this, we prove that the stage strategy sequence of both players are periodic after finite stages and the time-averaged strategy of player Y within one period is an exact NE strategy. Accordingly, we propose an asymmetric paradigm for solving two-player zero-sum games. For the special game with rational payoff matrix and an interior NE, the paradigm can output the precise NE strategy; for any general games we prove that the time-averaged strategy can converge to an approximate NE. In comparison to the NE-solving method via Hedge self-play, this HBR paradigm exhibits faster computation/convergence, better stability and can attain precise NE convergence in most real cases.

math.DS

The Optimal Strategy against Hedge Algorithm in Repeated Games

This paper aims to solve the optimal strategy against a well-known adaptive algorithm, the Hedge algorithm, in a finitely repeated $2\times 2$ zero-sum game. In the literature, related theoretical results are very rare. To this end, we make the evolution analysis for the resulting dynamical game system and build the action recurrence relation based on the Bellman optimality equation. First, we define the state and the State Transition Triangle Graph (STTG); then, we prove that the game system will behave in a periodic-like way when the opponent adopts the myopic best response. Further, based on the myopic path and the recurrence relation between the optimal actions at time-adjacent states, we can solve the optimal strategy of the opponent, which is proved to be periodic on the time interval truncated by a tiny segment and has the same period as the myopic path. Results in this paper are rigorous and inspiring, and the method might help solve the optimal strategy for general games and general algorithms.

math.OC