Searcharxiv⌕ Search

arXiv subjects

Yuanhanqing Huang

Publications and source records attributed to Yuanhanqing Huang.

9 recordsLinked to original sources

Offline Learning of Decision Functions in Multiplayer Games with Expectation Constraints

We explore a class of stochastic multiplayer games where each player in the game aims to optimize its objective under uncertainty and adheres to some expectation constraints. The study employs an offline learning paradigm, leveraging a pre-existing dataset containing auxiliary features. While prior research in deterministic and stochastic multiplayer games primarily explored vector-valued decisions, this work departs by considering function-valued decisions that incorporate auxiliary features as input. We leverage the law of large deviations and degree theory to establish the almost sure convergence of the offline learning solution to the true solution as the number of data samples increases.

math.OC↗

Bandit Online Learning in Merely Coherent Games with Multi-Point Pseudo-Gradient Estimate

Non-cooperative games serve as a powerful framework for capturing the interactions among self-interested players and have broad applicability in modeling a wide range of practical scenarios, ranging from power management to drug delivery. Although most existing solution algorithms assume the availability of first-order information or full knowledge of the objectives and others' action profiles, there are situations where the only accessible information at players' disposal is the realized objective function values. In this paper, we devise a bandit online learning algorithm for merely coherent games that integrates the optimistic mirror descent scheme and multi-point pseudo-gradient estimates. We further demonstrate that the generated actual sequence of play can converge a.s. to a critical point if the sequences of query radius and sample size are chosen properly, without resorting to extra Tikhonov regularization terms or additional norm conditions. Finally, we illustrate the validity of the proposed algorithm via a Rock-Paper-Scissors game and a least square estimation game.

math.OC↗

A Bandit Learning Method for Continuous Games under Feedback Delays with Residual Pseudo-Gradient Estimate

Learning in multi-player games can model a large variety of practical scenarios, where each player seeks to optimize its own local objective function, which at the same time relies on the actions taken by others. Motivated by the frequent absence of first-order information such as partial gradients in solving local optimization problems and the prevalence of asynchronicity and feedback delays in multi-agent systems, we introduce a bandit learning algorithm, which integrates mirror descent, residual pseudo-gradient estimates, and the priority-based feedback utilization strategy, to contend with these challenges. We establish that for pseudo-monotone plus games, the actual sequences of play generated by the proposed algorithm converge a.s. to critical points. Compared with the existing method, the proposed algorithm yields more consistent estimates with less variation and allows for more aggressive choices of parameters. Finally, we illustrate the validity of the proposed algorithm through a thermal load management problem of building complexes.

math.OC↗

On the Convergence Rates of A Nash Equilibrium Seeking Algorithm in Potential Games with Information Delays

This paper investigates the equilibrium convergence properties of a proposed algorithm for potential games with continuous strategy spaces in the presence of feedback delays, a main challenge in multi-agent systems that compromises the performance of various optimization schemes. The proposed algorithm is built upon an improved version of the accelerated gradient descent method. We extend it to a decentralized multi-agent scenario and equip it with a delayed feedback utilization scheme. By appropriately tuning the step sizes and studying the interplay between delay functions and step sizes, we derive the convergence rates of the proposed algorithm to the optimal value of the potential function when the growth of the feedback delays in time is subject to sublinear, linear, and superlinear upper bounds. Finally, simulations of a routing game are performed to empirically verify our findings.

math.OC↗

Zeroth-Order Learning in Continuous Games via Residual Pseudogradient Estimates

A variety of practical problems can be modeled by the decision-making process in multi-player games where a group of self-interested players aim at optimizing their own local objectives, while the objectives depend on the actions taken by others. The local gradient information of each player, essential in implementing algorithms for finding game solutions, is all too often unavailable. In this paper, we focus on designing solution algorithms for multi-player games using bandit feedback, i.e., the only available feedback at each player's disposal is the realized objective values. To tackle the issue of large variances in the existing bandit learning algorithms with a single oracle call, we propose two algorithms by integrating the residual feedback scheme into single-call extra-gradient methods. Subsequently, we show that the actual sequences of play can converge almost surely to a critical point if the game is pseudo-monotone plus and characterize the convergence rate to the critical point when the game is strongly pseudo-monotone. The ergodic convergence rates of the generated sequences in monotone games are also investigated as a supplement. Finally, the validity of the proposed algorithms is further verified via numerical examples.

math.OC↗

Distributed Stochastic Nash Equilibrium Learning in Locally Coupled Network Games with Unknown Parameters

In stochastic Nash equilibrium problems (SNEPs), it is natural for players to be uncertain about their complex environments and have multi-dimensional unknown parameters in their models. Among various SNEPs, this paper focuses on locally coupled network games where the objective of each rational player is subject to the aggregate influence of its neighbors. We propose a distributed learning algorithm based on the proximal-point iteration and ordinary least-square estimator, where each player repeatedly updates the local estimates of neighboring decisions, makes its augmented best-response decisions given the current estimated parameters, receives the realized objective values, and learns the unknown parameters. Leveraging the Robbins-Siegmund theorem and the law of large deviations for M-estimators, we establish the almost sure convergence of the proposed algorithm to solutions of SNEPs when the updating step sizes decay at a proper rate.

math.OC↗

Distributed Computation of Stochastic GNE with Partial Information: An Augmented Best-Response Approach

In this paper, we focus on the stochastic generalized Nash equilibrium problem (SGNEP) which is an important and widely-used model in many different fields. In this model, subject to certain global resource constraints, a set of self-interested players aim to optimize their local objectives that depend on their own decisions and the decisions of others and are influenced by some random factors. We propose a distributed stochastic generalized Nash equilibrium seeking algorithm in a partial-decision information setting based on the Douglas-Rachford operator splitting scheme, which relaxes assumptions in the existing literature. The proposed algorithm updates players' local decisions through augmented best-response schemes and subsequent projections onto the local feasible sets, which occupy most of the computational workload. The projected stochastic subgradient method is applied to provide approximate solutions to the augmented best-response subproblems for each player. The Robbins-Siegmund theorem is leveraged to establish the main convergence results to a true Nash equilibrium using the proposed inexact solver. Finally, we illustrate the validity of the proposed algorithm via two numerical examples, i.e., a stochastic Nash-Cournot distribution game and a multi-product assembly problem with the two-stage model.

eess.SY↗

A Primal Decomposition Approach to Globally Coupled Aggregative Optimization over Networks

We consider a class of multi-agent optimization problems, where each agent has a local objective function that depends on its own decision variables and the aggregate of others, and is willing to cooperate with other agents to minimize the sum of the local objectives. After associating each agent with an auxiliary variable and the related local estimates, we conduct primal decomposition to the globally coupled problem and reformulate it so that it can be solved distributedly. Based on the Douglas-Rachford method, an algorithm is proposed which ensures the exact convergence to a solution of the original problem. The proposed method enjoys desirable scalability by only requiring each agent to keep local estimates whose number grows linearly with the number of its neighbors. We illustrate our proposed algorithm by numerical simulations on a commodity distribution problem over a transport network.

eess.SY↗

A Distributed GNE Seeking Algorithm Using the Douglas-Rachford Splitting Method

We consider a generalized Nash equilibrium problem (GNEP) for a network of players. Each player tries to minimize a local objective function subject to some resource constraints where both the objective functions and the resource constraints depend on other players' decisions. By conducting equivalent transformations on the local optimization problems and introducing network Lagrangian, we recast the GNEP into an operator zero-finding problem. An algorithm is proposed based on the Douglas-Rachford method to distributedly find a solution. The proposed algorithm requires milder conditions compared to the existing methods. We prove the convergence of the proposed algorithm to an exact variational generalized Nash equilibrium under two different sets of assumptions. Our algorithm is validated numerically through the example of a Nash-Cournot production game.

math.OC↗