SearcharxivSearch

arXiv subjects

Sarper Aydin

Publications and source records attributed to Sarper Aydin.

3 recordsLinked to original sources

Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

We propose networked policy gradient play for solving Markov potential games with continuous and/or discrete state-action pairs. During the game, agents use parametrized and differentiable policies that depend on the current state and the policy parameters of other agents. During training, agents update their policy parameters following stochastic gradients. The gradient estimation involves two consecutive episodes, generating unbiased estimators of reward and policy score functions. In addition, it involves keeping estimates of others' parameters using consensus steps given local estimates received through a time-varying communication network. In Markov potential games, there exists a potential value function among agents with gradients corresponding to the gradients of local value functions. Using this structure, we prove almost sure convergence to a stationary point of the potential value function with rate $O(1/ε^2)$. Compared to previous works, our results do not require bounded policy gradients or initial agreement on the values of individual policy parameters. Numerical experiments on a dynamic multi-agent newsvendor problem verify the convergence of local beliefs and gradients. It further shows that networked policy gradient play converges as fast as independent policy gradient updates, while collecting higher rewards.

eess.SY

Decentralized Fictitious Play Converges Near a Nash Equilibrium in Near-Potential Games

We investigate convergence of decentralized fictitious play (DFP) in near-potential games, wherein agents preferences can almost be captured by a potential function. In DFP agents keep local estimates of other agents' empirical frequencies, best-respond against these estimates, and receive information over a time-varying communication network. We prove that empirical frequencies of actions generated by DFP converge around a single Nash Equilibrium (NE) assuming that there are only finitely many Nash equilibria, and the difference in utility functions resulting from unilateral deviations is close enough to the difference in the potential function values. This result assures that DFP has the same convergence properties of standard Fictitious play (FP) in near-potential games.

cs.GT

A Best-Response Algorithm with Voluntary Communication and Mobility Protocols for Mobile Autonomous Teams Solving the Target Assignment Problem

We consider a team of mobile autonomous robots with the aim to cover a given set of targets. Each robot aims to select a target to cover and physically reach it by the final time in coordination with other robots given the locations of targets. Robots are unaware of which targets other robots intend to cover. Each robot can control its mobility and who to send information to. We assume communication happens over a wireless channel that is subject to fading and failures. Given the setup, we propose a decentralized algorithm based on decentralized fictitious play in which robots reason about the selections and locations of other robots to decide which target to select, whether to communicate or not, who to communicate with, and where to move. Specifically, the communication actions of the robots are learning-aware, and their mobility actions are sensitive to the success probability of communication. We show that the decentralized algorithm guarantees that robots will cover their targets in finite time. Numerical simulations and experiments using a team of mobile robots confirm the target coverage in finite time and show that mobility control for communication and learning-aware voluntary communication protocols reduce the number of communication attempts in comparison to a benchmark distributed algorithm that relies on communication after every decision epoch.

eess.SY