SearcharxivSearch

arXiv subjects

Lacra Pavel

Publications and source records attributed to Lacra Pavel.

At least 19 recordsLinked to original sources

Safe Distributed Generalized Nash Equilibrium Seeking via Control Barrier Functions

In this paper, we consider generalized Nash equilibrium (GNE) seeking in non-cooperative games with coupled constraint sets. Specifically, we aim to enforce safety for distributed GNE seeking, whereby the safety specifications are encoded in the coupled constraint set. To achieve this, we introduce the control barrier function (CBF) in the design of the GNE seeking dynamics. We design the dynamics for both full- and partial-information setting, where each player has knowledge of the decision information of all other players or only neighboring players, respectively. We justify the proposed dynamics by showing that the coupled constraint set is forward invariant, the equilibrium of the dynamics coincides with the exact GNE of the game, and the dynamics is asymptotically stable. Furthermore, we extend the approach to games where the agents are multi-integrators. Numerical simulations are provided to verify our results.

math.OC

Satisficing Paths to Equilibrium, Generalized Weakly Acyclic Games, and Learning

Weakly acyclic games generalize potential games and have shown to be fundamental in the study of multi-agent learning as they allow for convergence to an equilibrium via best-responding under inertia. In this paper, we present a generalization of weakly acyclic games, and we demonstrate its importance in multi-agent learning when agents employ experimental strategy updates in periods where they fail to best respond. While weak acyclicity is defined in terms of path connectivity properties of a game's better response graph, our concept is defined using a generalized better response graph under revision dynamics termed as satisficing. We refer to this class of games as generalized weakly acyclic games (GenWAGs). We provide sufficient conditions for this notion of generalized weak acyclicity in both two-player games and n-player games in normal form, including static and dynamic games. Several graph theoretic characterizations of such games are presented together with sufficiency conditions, examples, and counterexamples. Finally, implications on learning via policy revision processes are presented.

cs.GT

Robust Accelerated Dynamics for Subnetwork Bilinear Zero-Sum Games with Distributed Restarting

In this paper, we investigate distributed Nash equilibrium seeking for a class of two-subnetwork zero-sum games characterized by bilinear coupling. We present a distributed primal-dual accelerated mirror-descent algorithm with convergence guarantees. However, we demonstrate that this time-varying algorithm is not robust, as it fails to converge under arbitrarily small disturbances. To address this limitation, we introduce a distributed accelerated algorithm that incorporates a coordinated restarting mechanism. We model this new algorithm as a hybrid dynamical system and establish its structural robustness.

math.OC

Primal-dual Accelerated Mirror-Descent Method for Constrained Bilinear Saddle-Point Problems

We develop a first-order accelerated algorithm for a class of constrained bilinear saddle-point problems with applications to network systems. The algorithm is a modified time-varying primal-dual version of an accelerated mirror-descent dynamics. It deals with constraints such as simplices and convex set constraints effectively, and converges with a rate of $O(1/t^2)$. Furthermore, we employ the acceleration scheme to constrained distributed optimization and bilinear zero-sum games, and obtain two variants of distributed accelerated algorithms.

math.OC

Paths to Equilibrium in Games

In multi-agent reinforcement learning (MARL) and game theory, agents repeatedly interact and revise their strategies as new data arrives, producing a sequence of strategy profiles. This paper studies sequences of strategies satisfying a pairwise constraint inspired by policy updating in reinforcement learning, where an agent who is best responding in one period does not switch its strategy in the next period. This constraint merely requires that optimizing agents do not switch strategies, but does not constrain the non-optimizing agents in any way, and thus allows for exploration. Sequences with this property are called satisficing paths, and arise naturally in many MARL algorithms. A fundamental question about strategic dynamics is such: for a given game and initial strategy profile, is it always possible to construct a satisficing path that terminates at an equilibrium? The resolution of this question has implications about the capabilities or limitations of a class of MARL algorithms. We answer this question in the affirmative for normal-form games. Our analysis reveals a counterintuitive insight that reward deteriorating strategic updates are key to driving play to equilibrium along a satisficing path.

cs.GT

Passivity-based Gradient-Play Dynamics for Distributed Generalized Nash Equilibrium Seeking

We consider seeking generalized Nash equilibria (GNE) for noncooperative games with coupled nonlinear constraints over networks. We first revisit a well-known gradientplay dynamics from a passivity-based perspective, and address that the strict monotonicity on pseudo-gradients is a critical assumption to ensure the exact convergence of the dynamics. Then we propose two novel passivity-based gradient-play dynamics by introducing parallel feedforward compensators (PFCs) and output feedback compensators (OFCs). We show that the proposed dynamics can reach exact GNEs in merely monotone regimes if the PFCs are strictly passive or the OFCs are output strictly passive. Following that, resorting to passivity, we develop a unifying framework to generalize the gradient-play dynamics, and moreover, design a class of explicit passive-based dynamics with convergence guarantees. In addition, we explore the relation between the proposed dynamics and some existing methods, and extend our results to partial-decision information settings.

math.OC

Second-Order Mirror Descent: Convergence in Games Beyond Averaging and Discounting

In this paper, we propose a second-order extension of the continuous-time game-theoretic mirror descent (MD) dynamics, referred to as MD2, which provably converges to mere (but not necessarily strict) variationally stable states (VSS) without using common auxiliary techniques such as time-averaging or discounting. We show that MD2 enjoys no-regret as well as an exponential rate of convergence towards strong VSS upon a slight modification. MD2 can also be used to derive many novel continuous-time primal-space dynamics. We then use stochastic approximation techniques to provide a convergence guarantee of discrete-time MD2 with noisy observations towards interior mere VSS. Selected simulations are provided to illustrate our results.

math.OC

Recursive Reasoning in Minimax Games: A Level $k$ Gradient Play Method

Despite the success of generative adversarial networks (GANs) in generating visually appealing images, they are notoriously challenging to train. In order to stabilize the learning dynamics in minimax games, we propose a novel recursive reasoning algorithm: Level $k$ Gradient Play (Lv.$k$ GP) algorithm. In contrast to many existing algorithms, our algorithm does not require sophisticated heuristics or curvature information. We show that as $k$ increases, Lv.$k$ GP converges asymptotically towards an accurate estimation of players' future strategy. Moreover, we justify that Lv.$\infty$ GP naturally generalizes a line of provably convergent game dynamics which rely on predictive updates. Furthermore, we provide its local convergence property in nonconvex-nonconcave zero-sum games and global convergence in bilinear and quadratic games. By combining Lv.$k$ GP with Adam optimizer, our algorithm shows a clear advantage in terms of performance and computational overhead compared to other methods. Using a single Nvidia RTX3090 GPU and 30 times fewer parameters than BigGAN on CIFAR-10, we achieve an FID of 10.17 for unconditional image generation within 30 hours, allowing GAN training on common computational resources to reach state-of-the-art performance.

cs.LG

Continuous-Time Convergence Rates in Potential and Monotone Games

In this paper, we provide exponential rates of convergence to the interior Nash equilibrium for continuous-time dual-space game dynamics such as mirror descent (MD) and actor-critic (AC). We perform our analysis in $N$-player continuous concave games that satisfy certain monotonicity assumptions while possibly also admitting potential functions. In the first part of this paper, we provide a novel relative characterization of monotone games and show that MD and its discounted version converge with $\mathcal{O}(e^{-βt})$ in relatively strongly and relatively hypo-monotone games, respectively. In the second part of this paper, we specialize our results to games that admit a relatively strongly concave potential and show AC converges with $\mathcal{O}(e^{-βt})$. These rates extend their known convergence conditions. Simulations are performed which empirically back up our results.

math.OC

Resilient Nash Equilibrium Seeking in the Partial Information Setting

Current research in distributed Nash equilibrium (NE) seeking in the partial information setting assumes that information is exchanged between agents that are "truthful". However, in general noncooperative games agents may consider sending misinformation to neighboring agents with the goal of further reducing their cost. Additionally, communication networks are vulnerable to attacks from agents outside the game as well as communication failures. In this paper, we propose a distributed NE seeking algorithm that is robust against adversarial agents that transmit noise, random signals, constant singles, deceitful messages, as well as being resilient to external factors such as dropped communication, jammed signals, and man in the middle attacks. The core issue that makes the problem challenging is that agents have no means of verifying if the information they receive is correct, i.e. there is no "ground truth". To address this problem, we use an observation graph, that gives truthful action information, in conjunction with a communication graph, that gives (potentially incorrect) information. By filtering information obtained from these two graphs, we show that our algorithm is resilient against adversarial agents and converges to the Nash equilibrium.

math.OC

An inexact-penalty method for GNE seeking in games with dynamic agents

We consider a network of autonomous agents whose outputs are actions in a game with coupled constraints. In such network scenarios, agents seeking to minimize coupled cost functions using distributed information while satisfying the coupled constraints. Current methods consider the small class of multi-integrator agents using primal-dual methods. These methods can only ensure constraint satisfaction in steady-state. In contrast, we propose an inexact penalty method using a barrier function for nonlinear agents with equilibrium-independent passive dynamics. We show that these dynamics converge to an epsilon-GNE while satisfying the constraints for all time, not only in steady-state. We develop these dynamics in both the full-information and partial-information settings. In the partial-information setting, dynamic estimates of the others' actions are used to make decisions and are updated through local communication. Applications to optical networks and velocity synchronization of flexible robots are provided.

eess.SY

On the exact convergence to Nash equilibrium in hypomonotone regimes under full and partial-information

In this paper, we consider distributed Nash equilibrium seeking in monotone and hypomonotone games. We first assume that each player has knowledge of the opponents' decisions and propose a passivity-based modification of the standard gradient-play dynamics, that we call "Heavy Anchor". We prove that Heavy Anchor allows a relaxation of strict monotonicity of the pseudo-gradient, needed for gradient-play dynamics, and can ensure exact asymptotic convergence in merely monotone regimes. We extend these results to the setting where each player has only partial information of the opponents' decisions. Each player maintains a local decision variable and an auxiliary state estimate and communicates with their neighbours to learn the opponents' actions. We modify Heavy Anchor via a distributed Laplacian feedback and show how we can exploit equilibrium-independent passivity properties to achieve convergence to a Nash equilibrium in hypomonotone regimes.

math.OC

Single-timescale distributed GNE seeking for aggregative games over networks via forward-backward operator splitting

We consider aggregative games with affine coupling constraints, where agents have partial information on the aggregate value and can only communicate with neighbouring agents. We propose a single-layer distributed algorithm that reaches a variational generalized Nash equilibrium, under constant step sizes. The algorithm works on a single timescale, i.e., does not require multiple communication rounds between agents before updating their action. The convergence proof leverages an invariance property of the aggregate estimates and relies on a forward-backward splitting for two preconditioned operators and their restricted (strong) monotonicity properties on the consensus subspace.

math.OC

Continuous-time Discounted Mirror-Descent Dynamics in Monotone Concave Games

In this paper, we consider concave continuous-kernel games characterized by monotonicity properties and propose discounted mirror descent-type dynamics. We introduce two classes of dynamics whereby the associated mirror map is constructed based on a strongly convex or a Legendre regularizer. Depending on the properties of the regularizer we show that these new dynamics can converge asymptotically in concave games with monotone (negative) pseudo-gradient. Furthermore, we show that when the regularizer enjoys strong convexity, the resulting dynamics can converge even in games with hypo-monotone (negative) pseudo-gradient, which corresponds to a shortage of monotonicity.

math.OC

On seeking efficient Pareto optimal points in multi-player minimum cost flow problems with application to transportation systems

In this paper, we propose a multi-player extension of the minimum cost flow problem inspired by a transportation problem that arises in modern transportation industry. We associate one player with each arc of a directed network, each trying to minimize its cost function subject to the network flow constraints. In our model, the cost function can be any general nonlinear function, and the flow through each arc is an integer. We present algorithms to compute efficient Pareto optimal point(s), where the maximum possible number of players (but not all) minimize their cost functions simultaneously. The computed Pareto optimal points are Nash equilibriums if the problem is transformed into a finite static game in normal form.

math.OC

Dynamic NE Seeking for Multi-Integrator Networked Agents with Disturbance Rejection

In this paper, we consider game problems played by (multi)-integrator agents, subject to external disturbances. We propose Nash equilibrium seeking dynamics based on gradient-play, augmented with a dynamic internal-model based component, which is a reduced-order observer of the disturbance. We consider single-, double- and extensions to multi-integrator agents, in a partial-information setting, where agents have only partial knowledge on the others' decisions over a network. The lack of global information is offset by each agent maintaining an estimate of the others' states, based on local communication with its neighbours. Each agent has an additional dynamic component that drives its estimates to the consensus subspace. In all cases, we show convergence to the Nash equilibrium irrespective of disturbances. Our proofs leverage input-to-state stability under strong monotonicity of the pseudo-gradient and Lipschitz continuity of the extended pseudo-gradient.

math.OC

Distributed Nash Equilibrium Seeking under Partial-Decision Information via the Alternating Direction Method of Multipliers

In this paper, we consider the problem of finding a Nash equilibrium in a multi-player game over generally connected networks. This model differs from a conventional setting in that players have partial information on the actions of their opponents and the communication graph is not necessarily the same as the players' cost dependency graph. We develop a relatively fast algorithm within the framework of inexact-ADMM, based on local information exchange between the players. We prove its convergence to Nash equilibrium for fixed step-sizes and analyze its convergence rate. Numerical simulations illustrate its benefits when compared to a consensus-based gradient type algorithm with diminishing step-sizes.

math.OC

From Game-theoretic Multi-agent Log Linear Learning to Reinforcement Learning

The main focus of this paper is on enhancement of two types of game-theoretic learning algorithms: log-linear learning and reinforcement learning. The standard analysis of log-linear learning needs a highly structured environment, i.e. strong assumptions about the game from an implementation perspective. In this paper, we introduce a variant of log-linear learning that provides asymptotic guarantees while relaxing the structural assumptions to include synchronous updates and limitations in information available to the players. On the other hand, model-free reinforcement learning is able to perform even under weaker assumptions on players' knowledge about the environment and other players' strategies. We propose a reinforcement algorithm that uses a double-aggregation scheme in order to deepen players' insight about the environment and constant learning step-size which achieves a higher convergence rate. Numerical experiments are conducted to verify each algorithm's robustness and performance.

cs.LG