Searcharxiv⌕ Search

arXiv subjects

Gürdal Arslan

Publications and source records attributed to Gürdal Arslan.

10 recordsLinked to original sources

Satisficing Paths to Equilibrium, Generalized Weakly Acyclic Games, and Learning

Weakly acyclic games generalize potential games and have shown to be fundamental in the study of multi-agent learning as they allow for convergence to an equilibrium via best-responding under inertia. In this paper, we present a generalization of weakly acyclic games, and we demonstrate its importance in multi-agent learning when agents employ experimental strategy updates in periods where they fail to best respond. While weak acyclicity is defined in terms of path connectivity properties of a game's better response graph, our concept is defined using a generalized better response graph under revision dynamics termed as satisficing. We refer to this class of games as generalized weakly acyclic games (GenWAGs). We provide sufficient conditions for this notion of generalized weak acyclicity in both two-player games and n-player games in normal form, including static and dynamic games. Several graph theoretic characterizations of such games are presented together with sufficiency conditions, examples, and counterexamples. Finally, implications on learning via policy revision processes are presented.

cs.GT↗

Fictitious Play in Extensive-Form Games of Imperfect Information

We study the long-term behavior of the fictitious play process in repeated extensive-form games of imperfect information with perfect recall. Each player maintains incorrect beliefs that the moves at all information sets, except the one at which the player is about to make a move, are made according to fixed random strategies, independently across all information sets. Accordingly, each player makes his moves at any of his information sets to maximize his expected payoff assuming that, at any other information set, the moves are made according to the empirical frequencies of the past moves. We extend the well-known Monderer-Shapley result [1] on the convergence of the empirical frequencies to the set of Nash equilibria to a certain class of extensive-form games with identical interests. We then strengthen this result by the use of inertia and fading memory, and prove the convergence of the realized play-paths to an essentially pure Nash equilibrium in all extensive-form games of imperfect information with identical interests.

cs.GT↗

Unsynchronized Decentralized Q-Learning: Two Timescale Analysis By Persistence

Non-stationarity is a fundamental challenge in multi-agent reinforcement learning (MARL), where agents update their behaviour as they learn. Many theoretical advances in MARL avoid the challenge of non-stationarity by coordinating the policy updates of agents in various ways, including synchronizing times at which agents are allowed to revise their policies. Synchronization enables analysis of many MARL algorithms via multi-timescale methods, but such synchronization is infeasible in many decentralized applications. In this paper, we study an unsynchronized variant of the decentralized Q-learning algorithm, a recent MARL algorithm for stochastic games. We provide sufficient conditions under which the unsynchronized algorithm drives play to equilibrium with high probability. Our solution utilizes constant learning rates in the Q-factor update, which we show to be critical for relaxing the synchronization assumptions of earlier work. Our analysis also applies to unsynchronized generalizations of a number of other algorithms from the regret testing tradition, whose performance is analyzed by multi-timescale methods that study Markov chains obtained via policy update dynamics. This work extends the applicability of the decentralized Q-learning algorithm and its relatives to settings in which parameters are selected in an independent manner, and tames non-stationarity without imposing the coordination assumptions of prior work.

cs.GT↗

Mean-Field Games With Finitely Many Players: Independent Learning and Subjectivity

Independent learners are agents that employ single-agent algorithms in multi-agent systems, intentionally ignoring the effect of other strategic agents. This paper studies mean-field games from a decentralized learning perspective, with two primary objectives: (i) to identify structure that can guide algorithm design, and (ii) to understand the emergent behaviour in systems of independent learners. We study a new model of partially observed mean-field games with finitely many players, local action observability, and a general observation channel for partial observations of the global state. Specific observation channels considered include (a) global observability, (b) local and mean-field observability, (c) local and compressed mean-field observability, and (d) only local observability. We establish conditions under which the control problem of a given agent is equivalent to a fully observed MDP, as well as conditions under which the control problem is equivalent only to a POMDP. Building on the connection to MDPs, we prove the existence of perfect equilibrium among memoryless stationary policies under mean-field observability. Leveraging the connection to POMDPs, we prove convergence of learning iterates obtained by independent learning agents under any of the aforementioned observation channels. We interpret the limiting values as subjective value functions, which an agent believes to be relevant to its control problem. These subjective value functions are then used to propose subjective Q-equilibrium, a new solution concept for partially observed n-player mean-field games, whose existence is proved under mean-field or global observability. We provide a decentralized learning algorithm for partially observed n-player mean-field games, and we show that it drives play to subjective Q-equilibrium by adapting the recently developed theory of satisficing paths to allow for subjectivity.

cs.GT↗

Paths to Equilibrium in Games

In multi-agent reinforcement learning (MARL) and game theory, agents repeatedly interact and revise their strategies as new data arrives, producing a sequence of strategy profiles. This paper studies sequences of strategies satisfying a pairwise constraint inspired by policy updating in reinforcement learning, where an agent who is best responding in one period does not switch its strategy in the next period. This constraint merely requires that optimizing agents do not switch strategies, but does not constrain the non-optimizing agents in any way, and thus allows for exploration. Sequences with this property are called satisficing paths, and arise naturally in many MARL algorithms. A fundamental question about strategic dynamics is such: for a given game and initial strategy profile, is it always possible to construct a satisficing path that terminates at an equilibrium? The resolution of this question has implications about the capabilities or limitations of a class of MARL algorithms. We answer this question in the affirmative for normal-form games. Our analysis reveals a counterintuitive insight that reward deteriorating strategic updates are key to driving play to equilibrium along a satisficing path.

cs.GT↗

Subjective Equilibria under Beliefs of Exogenous Uncertainty: Linear Quadratic Case

We consider a stochastic dynamic game where players have their own linear state dynamics and quadratic cost functions. Players are coupled through some environment variables, generated by another linear system driven by the states and decisions of all players. Each player observes his own states realized up to the current time as well as the past realizations of his own decisions and the environment variables. Each player (incorrectly) believes that the environment variables are generated by an independent exogenous stochastic process. In this setup, we study the notion of ``subjective equilibrium under beliefs of exogenous uncertainty (SEBEU)'' introduced in our recent work arXiv:2005.01640. At an SEBEU, each player's strategy is optimal with respect to his subjective belief; moreover, the objective probability distribution of the environment variables is consistent with players' subjective beliefs. We construct an SEBEU in pure strategies, where each player strategy is an affine function of his own state and his estimate of the system state.

math.OC↗

Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games

In multi-agent reinforcement learning (MARL), independent learners are those that do not observe the actions of other agents in the system. Due to the decentralization of information, it is challenging to design independent learners that drive play to equilibrium. This paper investigates the feasibility of using satisficing dynamics to guide independent learners to approximate equilibrium in stochastic games. For $ε\geq 0$, an $ε$-satisficing policy update rule is any rule that instructs the agent to not change its policy when it is $ε$-best-responding to the policies of the remaining players; $ε$-satisficing paths are defined to be sequences of joint policies obtained when each agent uses some $ε$-satisficing policy update rule to select its next policy. We establish structural results on the existence of $ε$-satisficing paths into $ε$-equilibrium in both symmetric $N$-player games and general stochastic games with two players. We then present an independent learning algorithm for $N$-player symmetric games and give high probability guarantees of convergence to $ε$-equilibrium under self-play. This guarantee is made using symmetry alone, leveraging the previously unexploited structure of $ε$-satisficing paths.

cs.GT↗

Subjective Equilibria under Beliefs of Exogenous Uncertainty for Dynamic Games

We present a subjective equilibrium notion (called "subjective equilibrium under beliefs of exogenous uncertainty (SEBEU)" for stochastic dynamic games in which each player chooses her decisions under the (incorrect) belief that a stochastic environment process driving the system is exogenous whereas in actuality this process is a solution of closed-loop dynamics affected by each individual player. Players observe past realizations of the environment variables and their local information. At equilibrium, if players are given the full distribution of the stochastic environment process as if it were an exogenous process, they would have no incentive to unilaterally deviate from their strategies. This notion thus generalizes what is known as the static price-taking behavior in prior literature to a stochastic and dynamic setup. We establish existence of SEBEU, study various properties and present explicit solutions. We obtain the $ε$-Nash equilibrium property of SEBEU when there are many players.

math.OC↗

Decentralized Learning for Optimality in Stochastic Dynamic Teams and Games with Local Control and Global State Information

Stochastic dynamic teams and games are rich models for decentralized systems and challenging testing grounds for multi-agent learning. Previous work that guaranteed team optimality assumed stateless dynamics, or an explicit coordination mechanism, or joint-control sharing. In this paper, we present an algorithm with guarantees of convergence to team optimal policies in teams and common interest games. The algorithm is a two-timescale method that uses a variant of Q-learning on the finer timescale to perform policy evaluation while exploring the policy space on the coarser timescale. Agents following this algorithm are "independent learners": they use only local controls, local cost realizations, and global state information, without access to controls of other agents. The results presented here are the first, to our knowledge, to give formal guarantees of convergence to team optimality using independent learners in stochastic dynamic teams and common interest games.

math.OC↗

Decentralized Q-Learning for Stochastic Teams and Games

There are only a few learning algorithms applicable to stochastic dynamic teams and games which generalize Markov decision processes to decentralized stochastic control problems involving possibly self-interested decision makers. Learning in games is generally difficult because of the non-stationary environment in which each decision maker aims to learn its optimal decisions with minimal information in the presence of the other decision makers who are also learning. In stochastic dynamic games, learning is more challenging because, while learning, the decision makers alter the state of the system and hence the future cost. In this paper, we present decentralized Q-learning algorithms for stochastic games, and study their convergence for the weakly acyclic case which includes team problems as an important special case. The algorithm is decentralized in that each decision maker has access to only its local information, the state information, and the local cost realizations; furthermore, it is completely oblivious to the presence of other decision makers. We show that these algorithms converge to equilibrium policies almost surely in large classes of stochastic games.

math.OC↗