Searcharxiv⌕ Search

arXiv subjects

Ilai Bistritz

Publications and source records attributed to Ilai Bistritz.

13 recordsLinked to original sources

Cordial Learning: Distributed Training with Correlated Data

We consider a distributed learning task with agents that have correlated data. Specifically, the label of an agent depends on the input of other agents for the same sample, and these inputs are also correlated. Correlated data is the reality when agents share the same environment. Existing decentralized methods, such as federated learning, ignore the structure of the problem and perform poorly on correlated data. On the other hand, centralized approaches are infeasible due to privacy and communication constraints. We introduce cordial (correlated and distributed) learning to address this gap by sharing only low-dimensional outputs between the agents while training local models to extract informative signals from peers. This distributed learning induces a game in which the loss function of each agent depends on the models of others. Assuming a linear model, we prove that cordial learning converges with probability one to a globally optimal solution, despite the nonconvex global objective. Experiments on structured multi-digit MNIST tasks demonstrate that cordial learning remains highly effective even in highly nonlinear settings.

cs.LG↗

Choose Your Battles: Distributed Learning Over Multiple Tug of War Games

Consider $N$ players and $K$ games taking place simultaneously. Each of these games is modeled as a Tug-of-War (ToW) game where increasing the action of one player decreases the reward for all other players. Each player participates in only one game at any given time. At each time step, a player decides the game in which they wish to participate in and the action they take in that game. Their reward depends on the actions of all players that are in the same game. This system of $K$ games is termed a 'Meta Tug-of-War' (Meta-ToW) game. These games can model scenarios such as power control, distributed task allocation, and activation in sensor networks. We propose the Meta Tug-of-Peace algorithm, a distributed algorithm where the action updates are done using a simple stochastic approximation algorithm, and the decision to switch games is made using an infrequent 1-bit communication between the players. We prove that in Meta-ToW games, our algorithm converges to an equilibrium that satisfies a target Quality of Service reward vector for the players. We then demonstrate the efficacy of our algorithm through simulations for the scenarios mentioned above.

cs.GT↗

Learning to Control Unknown Strongly Monotone Games

Consider a strongly monotone game where the players' utility functions include a reward function and a linear term for each dimension, with coefficients that are controlled by the manager. Gradient play converges to a unique Nash equilibrium (NE) that does not optimize the global objective. The global performance at NE can be improved by imposing linear constraints on the NE, also known as a generalized Nash equilibrium (GNE). We therefore want the manager to control the coefficients such that they impose the desired constraint on the NE. However, this requires knowing the players' rewards and action sets. Obtaining this game information is infeasible in a large-scale network and violates user privacy. To overcome this, we propose a simple algorithm that learns to shift the NE to meet the linear constraints by adjusting the controlled coefficients online. Our algorithm only requires the linear constraints violation as feedback and does not need to know the reward functions or the action sets. We prove that our algorithm converges with probability 1 to the set of GNE given by coupled linear constraints. We then prove an L2 convergence rate of near-$O(t^{-1/4})$.

cs.MA↗

Equilibrium Bandits: Learning Optimal Equilibria of Unknown Dynamics

Consider a decision-maker that can pick one out of $K$ actions to control an unknown system, for $T$ turns. The actions are interpreted as different configurations or policies. Holding the same action fixed, the system asymptotically converges to a unique equilibrium, as a function of this action. The dynamics of the system are unknown to the decision-maker, which can only observe a noisy reward at the end of every turn. The decision-maker wants to maximize its accumulated reward over the $T$ turns. Learning what equilibria are better results in higher rewards, but waiting for the system to converge to equilibrium costs valuable time. Existing bandit algorithms, either stochastic or adversarial, achieve linear (trivial) regret for this problem. We present a novel algorithm, termed Upper Equilibrium Concentration Bound (UECB), that knows to switch an action quickly if it is not worth it to wait until the equilibrium is reached. This is enabled by employing convergence bounds to determine how far the system is from equilibrium. We prove that UECB achieves a regret of $\mathcal{O}(\log(T)+τ_c\log(τ_c)+τ_c\log\log(T))$ for this equilibrium bandit problem where $τ_c$ is the worst case approximate convergence time to equilibrium. We then show that both epidemic control and game control are special cases of equilibrium bandits, where $τ_c\log τ_c$ typically dominates the regret. We then test UECB numerically for both of these applications.

cs.LG↗

Do Informational Cascades Happen with Non-myopic Agents?

We consider an environment where players need to decide whether to buy a certain product (or adopt a technology) or not. The product is either good or bad, but its true value is unknown to the players. Instead, each player has her own private information on its quality. Each player can observe the previous actions of other players and estimate the quality of the product. A classic result in the literature shows that in similar settings informational cascades occur where learning stops for the whole network and players repeat the actions of their predecessors. In contrast to this literature, in this work, players get more than one opportunity to act. In each turn, a player is chosen uniformly at random from all players and can decide to buy the product and leave the market or wait. Her utility is the total expected discounted reward, and thus myopic strategies may not constitute equilibria. We provide a characterization of perfect Bayesian equilibria (PBE) with forward-looking strategies through a fixed-point equation of dimensionality that grows only quadratically with the number of players. Using this tractable fixed-point equation, we show the existence of a PBE and characterize PBE with threshold strategies. Based on this characterization we study informational cascades in two regimes. First, we show that for a discount factor δ strictly smaller than one, informational cascades happen with high probability as the number of players N increases. Furthermore, only a small portion of the total information in the system is revealed before a cascade occurs ...

econ.GN↗

No Weighted-Regret Learning in Adversarial Bandits with Delays

Consider a scenario where a player chooses an action in each round $t$ out of $T$ rounds and observes the incurred cost after a delay of $d_{t}$ rounds. The cost functions and the delay sequence are chosen by an adversary. We show that in a non-cooperative game, the expected weighted ergodic distribution of play converges to the set of coarse correlated equilibria if players use algorithms that have "no weighted-regret" in the above scenario, even if they have linear regret due to too large delays. For a two-player zero-sum game, we show that no weighted-regret is sufficient for the weighted ergodic average of play to converge to the set of Nash equilibria. We prove that the FKM algorithm with $n$ dimensions achieves an expected regret of $O\left(nT^{\frac{3}{4}}+\sqrt{n}T^{\frac{1}{3}}D^{\frac{1}{3}}\right)$ and the EXP3 algorithm with $K$ arms achieves an expected regret of $O\left(\sqrt{\log K\left(KT+D\right)}\right)$ even when $D=\sum_{t=1}^{T}d_{t}$ and $T$ are unknown. These bounds use a novel doubling trick that, under mild assumptions, provably retains the regret bound for when $D$ and $T$ are known. Using these bounds, we show that FKM and EXP3 have no weighted-regret even for $d_{t}=O\left(t\log t\right)$. Therefore, algorithms with no weighted-regret can be used to approximate a CCE of a finite or convex unknown game that can only be simulated with bandit feedback, even if the simulation involves significant delays.

cs.LG↗

Game of Thrones: Fully Distributed Learning for Multi-Player Bandits

We consider an N-player multi-armed bandit game where each player chooses one out of M arms for T turns. Each player has different expected rewards for the arms, and the instantaneous rewards are independent and identically distributed or Markovian. When two or more players choose the same arm, they all receive zero reward. Performance is measured using the expected sum of regrets, compared with an optimal assignment of arms to players that maximizes the sum of expected rewards. We assume that each player only knows her actions and the reward she received each turn. Players cannot observe the actions of other players, and no communication between players is possible. We present a distributed algorithm and prove that it achieves an expected sum of regrets of near-O\left(\log T\right). This is the first algorithm to achieve a near order optimal regret in this fully distributed scenario. All other works have assumed that either all players have the same vector of expected rewards or that communication between players is possible.

cs.GT↗

My Fair Bandit: Distributed Learning of Max-Min Fairness with Multi-player Bandits

Consider N cooperative but non-communicating players where each plays one out of M arms for T turns. Players have different utilities for each arm, representable as an NxM matrix. These utilities are unknown to the players. In each turn players select an arm and receive a noisy observation of their utility for it. However, if any other players selected the same arm that turn, all colliding players will all receive zero utility due to the conflict. No other communication or coordination between the players is possible. Our goal is to design a distributed algorithm that learns the matching between players and arms that achieves max-min fairness while minimizing the regret. We present an algorithm and prove that it is regret optimal up to a $\log\log T$ factor. This is the first max-min fairness multi-player bandit algorithm with (near) order optimal regret.

cs.GT↗

Distributed Learning for Channel Allocation Over a Shared Spectrum

Channel allocation is the task of assigning channels to users such that some objective (e.g., sum-rate) is maximized. In centralized networks such as cellular networks, this task is carried by the base station which gathers the channel state information (CSI) from the users and computes the optimal solution. In distributed networks such as ad-hoc and device-to-device (D2D) networks, no base station exists and conveying global CSI between users is costly or simply impractical. When the CSI is time varying and unknown to the users, the users face the challenge of both learning the channel statistics online and converge to a good channel allocation. This introduces a multi-armed bandit (MAB) scenario with multiple decision makers. If two users or more choose the same channel, a collision occurs and they all receive zero reward. We propose a distributed channel allocation algorithm that each user runs and converges to the optimal allocation while achieving an order optimal regret of O\left(\log T\right). The algorithm is based on a carrier sensing multiple access (CSMA) implementation of the distributed auction algorithm. It does not require any exchange of information between users. Users need only to observe a single channel at a time and sense if there is a transmission on that channel, without decoding the transmissions or identifying the transmitting users. We demonstrate the performance of our algorithm using simulated LTE and 5G channels.

cs.IT↗

Approximate Best-Response Dynamics in Random Interference Games

In this paper we develop a novel approach to the convergence of Best-Response Dynamics for the family of interference games. Interference games represent the fundamental resource allocation conflict between users of the radio spectrum. In contrast to congestion games, interference games are generally not potential games. Therefore, proving the convergence of the best-response dynamics to a Nash equilibrium in these games requires new techniques. We suggest a model for random interference games, based on the long term fading governed by the players' geometry. Our goal is to prove convergence of the approximate best-response dynamics with high probability with respect to the randomized game. We embrace the asynchronous model in which the acting player is chosen at each stage at random. In our approximate best-response dynamics, the action of a deviating player is chosen at random among all the approximately best ones. We show that with high probability, with respect to the players' geometry and asymptotically with the number of players, each action increases the expected social-welfare (sum of achievable rates). Hence, the induced sum-rate process is a submartingale. Based on the Martingale Convergence Theorem, we prove convergence of the strategy profile to an approximate Nash equilibrium with good performance for asymptotically almost all interference games. We use the Markovity of the induced sum-rate process to provide probabilistic bounds on the convergence time. Finally, we demonstrate our results in simulated examples.

cs.GT↗

Game Theoretic Dynamic Channel Allocation for Frequency-Selective Interference Channels

We consider the problem of distributed channel allocation in large networks under the frequency-selective interference channel. Performance is measured by the weighted sum of achievable rates. Our proposed algorithm is a modified Fictitious Play algorithm that can be implemented distributedly and its stable points are the pure Nash equilibria of a given game. Our goal is to design a utility function for a non-cooperative game such that all of its pure Nash equilibria have close to optimal global performance. This will make the algorithm close to optimal while requiring no communication between users. We propose a novel technique to analyze the Nash equilibria of a random interference game, determined by the random channel gains. Our analysis is asymptotic in the number of users. First we present a natural non-cooperative game where the utility of each user is his achievable rate. It is shown that, asymptotically in the number of users and for strong enough interference, this game exhibits many bad equilibria. Then we propose a novel non-cooperative M Frequency-Selective Interference Channel Game (M-FSIG), as a slight modification of the former, where the utility of each user is artificially limited. We prove that even its worst equilibrium has asymptotically optimal weighted sum-rate for any interference regime and even for correlated channels. This is based on an order statistics analysis of the fading channels that is valid for a broad class of fading distributions (including Rayleigh, Rician, m-Nakagami and more). We carry out simulations that show fast convergence of our algorithm to the proven asymptotically optimal pure Nash equilibria.

cs.IT↗

Asymptotically Optimal Resource Block Allocation With Limited Feedback

Consider a channel allocation problem over a frequency-selective channel.There are K channels (frequency bands) and N users such that K=bN for some positive integer b. We want to allocate b channels (or resource blocks) to each user. Due to the nature of the frequency-selective channel, each user considers some channels to be better than others. The optimal solution to this resource allocation problem can be computed using the Hungarian algorithm. However, this requires knowledge of the numerical value of all the channel gains, which makes this approach impractical for large networks. We suggest a suboptimal approach, that only requires knowing what the M-best channels of each user are. We find the minimal value of M such that there exists an allocation where all the b channels each user gets are among his M-best. This leads to feedback of significantly less than one bit per user per channel. For a large class of fading distributions, including Rayleigh, Rician, m-Nakagami and others, this suboptimal approach leads to both an asymptotically (in K) optimal sum-rate and an asymptotically optimal minimal rate. Our non-opportunistic approach achieves (asymptotically) full multiuser diversity as well as optimal fairness, by contrast to all other limited feedback algorithms.

cs.IT↗

Asymptotically Optimal Distributed Channel Allocation: a Competitive Game-Theoretic Approach

In this paper we consider the problem of distributed channel allocation in large networks under the frequency-selective interference channel. Performance is measured by the weighted sum of achievable rates. First we present a natural non-cooperative game theoretic formulation for this problem. It is shown that, when interference is sufficiently strong, this game has a pure price of anarchy approaching infinity with high probability, and there is an asymptotically increasing number of equilibria with the worst performance. Then we propose a novel non-cooperative M Frequency-Selective Interference Game (M-FSIG), where users limit their utility such that it is greater than zero only for their M best channels, and equal for them. We show that the M-FSIG exhibits, with high probability, an increasing number of optimal pure Nash equilibria and no bad equilibria. Consequently, the pure price of anarchy converges to one in probability in any interference regime. In order to exploit these results algorithmically we propose a modified Fictitious Play algorithm that can be implemented distributedly. We carry out simulations that show its fast convergence to the proven pure Nash equilibria.

cs.GT↗