SearcharxivSearch

arXiv subjects

Yuma Fujimoto

Publications and source records attributed to Yuma Fujimoto.

18 recordsLinked to original sources

Immune history shapes recurrent epidemics of antigenically related variants

Population immunity carried over from past epidemics of an antigenically variable pathogen influences the epidemic of new variants based on their antigenic similarity to the previous ones. We develop a recurrent SIR model where a population faces sequential, antigenically related variants. The model yields a recurrence map for the population susceptibility to successive variants under the assumption of status-based population immunity. The model reveals that stable, equal-sized recurrent epidemics occur across broad parameter ranges, but can be destabilized when transmission is strong and antigenic escape is limited, leading to period-2 or more, or even more complex epidemic dynamics. Epidemic size is maximized at an intermediate basic reproduction number: higher transmissibility boosts immediate infection but also enhances cross-immunity, reducing future susceptibility of the population. Our results clarify how immune history shapes recurrent epidemics and why success in one wave does not ensure larger future epidemics.

q-bio.PE

Linear Convergence in Games with Delayed Feedback via Extra Prediction

Feedback delays are inevitable in real-world multi-agent learning. They are known to severely degrade performance, and the convergence rate under delayed feedback is still unclear, even for bilinear games. This paper derives the rate of linear convergence of Weighted Optimistic Gradient Descent-Ascent (WOGDA), which predicts future rewards with extra optimism, in unconstrained bilinear games. To analyze the algorithm, we interpret it as an approximation of the Extra Proximal Point (EPP), which is updated based on farther future rewards than the classical Proximal Point (PP). Our theorems show that standard optimism (predicting the next-step reward) achieves linear convergence to the equilibrium at a rate $\exp(-\Theta(t/m^{5}))$ after $t$ iterations for delay $m$. Moreover, employing extra optimism (predicting farther future reward) tolerates a larger step size and significantly accelerates the rate to $\exp(-\Theta(t/(m^{2}\log m)))$. Our experiments also show accelerated convergence driven by the extra optimism and are qualitatively consistent with our theorems. In summary, this paper validates that extra optimism is a promising countermeasure against performance degradation caused by feedback delays.

cs.LG

Learning from Delayed Feedback in Games via Extra Prediction

This study raises and addresses the problem of time-delayed feedback in learning in games. Because learning in games assumes that multiple agents independently learn their strategies, a discrepancy in optimization often emerges among the agents. To overcome this discrepancy, the prediction of the future reward is incorporated into algorithms, typically known as Optimistic Follow-the-Regularized-Leader (OFTRL). However, the time delay in observing the past rewards hinders the prediction. Indeed, this study firstly proves that even a single-step delay worsens the performance of OFTRL from the aspects of social regret and convergence. This study proposes the weighted OFTRL (WOFTRL), where the prediction vector of the next reward in OFTRL is weighted $n$ times. We further capture an intuition that the optimistic weight cancels out this time delay. We prove that when the optimistic weight exceeds the time delay, our WOFTRL recovers the good performances that social regret is constant in general-sum normal-form games, and the strategies last-iterate converge to the Nash equilibrium in poly-matrix zero-sum games. The theoretical results are supported and strengthened by our experiments.

cs.LG

Hamiltonian of polymatrix zero-sum games

The understanding of a dynamical system's properties can be significantly advanced by establishing it as a Hamiltonian system and then systematically exploring its inherent symmetries. By formulating agents' strategies and cumulative payoffs as canonically conjugate variables, we identify the Hamiltonian function that generates the dynamics of poly-matrix zero-sum games. We reveal the symmetries of our Hamiltonian and derive the associated conserved quantities, showing how the conservation of probability and the invariance of the Fenchel coupling are intrinsically encoded within the system. Furthermore, we propose the dissipation FTRL (DFTRL) dynamics by introducing a perturbation that dissipates the Fenchel coupling, proving convergence to the Nash equilibrium and linking DFTRL to last-iterate convergent algorithms. Our results highlight the potential of Hamiltonian dynamics in uncovering the structural properties of learning dynamics in games, and pave the way for broader applications of Hamiltonian dynamics in game theory and machine learning.

cs.GT

Time-Varyingness in Auction Breaks Revenue Equivalence

Auction is applied for trade with various mechanisms. A simple but practical question is which mechanism, typically first-price or second-price auctions, is preferred from the perspective of bidders or sellers. A celebrated answer is revenue equivalence, where each bidder's equilibrium payoff is proven to be independent of auction mechanisms (and a seller's revenue, too). In reality, however, auction environments like the value distribution of items would vary over time, and such equilibrium bidding cannot always be achieved. Indeed, bidders must continue to track their equilibrium bidding by learning in first-price auctions, but they can keep their equilibrium bidding in second-price auctions. This study discusses whether and how revenue equivalence is violated in the long run by comparing the time series of non-equilibrium bidding in first-price auctions with those of equilibrium bidding in second-price auctions. We characterize the value distribution by two parameters: its basis value, which means the lowest price to bid, and its value interval, which means the width of possible values. Surprisingly, our theorems and experiments find that revenue equivalence is broken by the correlation between the basis value and the value interval, uncovering a novel phenomenon that could occur in the real world.

cs.GT

Synchronization in Learning in Periodic Zero-Sum Games Triggers Divergence from Nash Equilibrium

Learning in zero-sum games studies a situation where multiple agents competitively learn their strategy. In such multi-agent learning, we often see that the strategies cycle around their optimum, i.e., Nash equilibrium. When a game periodically varies (called a ``periodic'' game), however, the Nash equilibrium moves generically. How learning dynamics behave in such periodic games is of interest but still unclear. Interestingly, we discover that the behavior is highly dependent on the relationship between the two speeds at which the game changes and at which players learn. We observe that when these two speeds synchronize, the learning dynamics diverge, and their time-average does not converge. Otherwise, the learning dynamics draw complicated cycles, but their time-average converges. Under some assumptions introduced for the dynamical systems analysis, we prove that this behavior occurs. Furthermore, our experiments observe this behavior even if removing these assumptions. This study discovers a novel phenomenon, i.e., synchronization, and gains insight widely applicable to learning in periodic games.

cs.GT

Global Behavior of Learning Dynamics in Zero-Sum Games with Memory Asymmetry

This study examines the global behavior of dynamics in learning in games between two players, X and Y. We consider the simplest situation for memory asymmetry between two players: X memorizes the other Y's previous action and uses reactive strategies, while Y has no memory. Although this memory complicates their learning dynamics, we characterize the global behavior of such complex dynamics by discovering and analyzing two novel quantities. One is an extended Kullback-Leibler divergence from the Nash equilibrium, a well-known conserved quantity from previous studies. The other is a family of Lyapunov functions of X's reactive strategy. One of the global behaviors we capture is that if X exploits Y, then their strategies converge to the Nash equilibrium. Another is that if Y's strategy is out of equilibrium, then X becomes more exploitative with time. Consequently, we suggest global convergence to the Nash equilibrium from both aspects of theory and experiment. This study provides a novel characterization of the global behavior in learning in games through a couple of indicators.

cs.GT

Memory Asymmetry Creates Heteroclinic Orbits to Nash Equilibrium in Learning in Zero-Sum Games

Learning in games considers how multiple agents maximize their own rewards through repeated games. Memory, an ability that an agent changes his/her action depending on the history of actions in previous games, is often introduced into learning to explore more clever strategies and discuss the decision-making of real agents like humans. However, such games with memory are hard to analyze because they exhibit complex phenomena like chaotic dynamics or divergence from Nash equilibrium. In particular, how asymmetry in memory capacities between agents affects learning in games is still unclear. In response, this study formulates a gradient ascent algorithm in games with asymmetry memory capacities. To obtain theoretical insights into learning dynamics, we first consider a simple case of zero-sum games. We observe complex behavior, where learning dynamics draw a heteroclinic connection from unstable fixed points to stable ones. Despite this complexity, we analyze learning dynamics and prove local convergence to these stable fixed points, i.e., the Nash equilibria. We identify the mechanism driving this convergence: an agent with a longer memory learns to exploit the other, which in turn endows the other's utility function with strict concavity. We further numerically observe such convergence in various initial strategies, action numbers, and memory lengths. This study reveals a novel phenomenon due to memory asymmetry, providing fundamental strides in learning in games and new insights into computing equilibria.

cs.GT

Nash Equilibrium and Learning Dynamics in Three-Player Matching $m$-Action Games

Learning in games discusses the processes where multiple players learn their optimal strategies through the repetition of game plays. The dynamics of learning between two players in zero-sum games, such as Matching Pennies, where their benefits are competitive, have already been well analyzed. However, it is still unexplored and challenging to analyze the dynamics of learning among three players. In this study, we formulate a minimalistic game where three players compete to match their actions with one another. Although interaction among three players diversifies and complicates the Nash equilibria, we fully analyze the equilibria. We also discuss the dynamics of learning based on some famous algorithms categorized into Follow the Regularized Leader. From both theoretical and experimental aspects, we characterize the dynamics by categorizing three-player interactions into three forces to synchronize their actions, switch their actions rotationally, and seek competition.

cs.GT

Game-theoretical approach to minimum entropy productions in information thermodynamics

In a situation where each player has control over the transition probabilities of each subsystem, we game-theoretically analyze the optimization problem of minimizing both the partial entropy production of each subsystem and a penalty for failing to achieve a given state transition. In the regime of linear irreversible thermodynamics, we obtain the Nash equilibrium solution of the probability flow and calculate each partial entropy production for this solution. We find a trade-off such that a partial entropy production should be larger if we want the other partial entropy production to be smaller. The total entropy production can be minimized if each subsystem equally shares the penalty. We identify that this trade-off is due to the interacting contribution of the probability flow and discuss a possible biological validity for Escherichia coli chemotaxis.

cond-mat.stat-mech

Evolutionary stability of cooperation by the leading eight norms in indirect reciprocity under noisy and private assessment

Indirect reciprocity is a mechanism that explains large-scale cooperation in human societies. In indirect reciprocity, an individual chooses whether or not to cooperate with another based on reputation information, and others evaluate the action as good or bad. Under what evaluation rule (called ``social norm'') cooperation evolves has long been of central interest in the literature. It has been reported that if individuals can share their evaluations (i.e., public reputation), social norms called ``leading eight'' can be evolutionarily stable. On the other hand, when they cannot share their evaluations (i.e., private assessment), the evolutionary stability of cooperation is still in question. To tackle this problem, we create a novel method to analyze the reputation structure in the population under private assessment. Specifically, we characterize each individual by two variables, ``goodness'' (what proportion of the population considers the individual as good) and ``self-reputation'' (whether an individual thinks of him/herself as good or bad), and analyze the stochastic process of how these two variables change over time. We discuss evolutionary stability of each of the leading eight social norms by studying the robustness against invasions of unconditional cooperators and defectors. We identify key pivots in those social norms for establishing a high level of cooperation or stable cooperation against mutants. Our finding gives an insight into how human cooperation is established in a real-world society.

q-bio.PE

Learning in Multi-Memory Games Triggers Complex Dynamics Diverging from Nash Equilibrium

Repeated games consider a situation where multiple agents are motivated by their independent rewards throughout learning. In general, the dynamics of their learning become complex. Especially when their rewards compete with each other like zero-sum games, the dynamics often do not converge to their optimum, i.e., the Nash equilibrium. To tackle such complexity, many studies have understood various learning algorithms as dynamical systems and discovered qualitative insights among the algorithms. However, such studies have yet to handle multi-memory games (where agents can memorize actions they played in the past and choose their actions based on their memories), even though memorization plays a pivotal role in artificial intelligence and interpersonal relationship. This study extends two major learning algorithms in games, i.e., replicator dynamics and gradient ascent, into multi-memory games. Then, we prove their dynamics are identical. Furthermore, theoretically and experimentally, we clarify that the learning dynamics diverge from the Nash equilibrium in multi-memory zero-sum games and reach heteroclinic cycles (sojourn longer around the boundary of the strategy space), providing a fundamental advance in learning in games.

cs.GT

Evolutionary stability of cooperation in indirect reciprocity under noisy and private assessment

Indirect reciprocity is a mechanism that explains large-scale cooperation in humans. In indirect reciprocity, individuals use reputations to choose whether or not to cooperate with a partner and update others' reputations. A major question is how the rules to choose their actions and the rules to update reputations evolve. In the public reputation case, where all individuals share the evaluation of others, social norms called Simple Standing (SS) and Stern Judging (SJ) have been known to maintain cooperation. However, in the case of private assessment where individuals independently evaluate others, the mechanism of maintenance of cooperation is still largely unknown. This study theoretically shows for the first time that cooperation by indirect reciprocity can be evolutionarily stable under private assessment. Specifically, we find that SS can be stable, but SJ can never be. This is intuitive because SS can correct interpersonal discrepancies in reputations through its simplicity. On the other hand, SJ is too complicated to avoid an accumulation of errors, which leads to the collapse of cooperation. We conclude that moderate simplicity is a key to success in maintaining cooperation under the private assessment. Our result provides a theoretical basis for evolution of human cooperation.

cs.GT

Reputation structure in indirect reciprocity under noisy and private assessment

Evaluation relationships are pivotal for maintaining a cooperative society. A formation of the evaluation relationships has been discussed in terms of indirect reciprocity, by modeling dynamics of good or bad reputations among individuals. Recently, a situation that individuals independently evaluate others with errors (i.e., noisy and private reputation) is considered, where the reputation structure (from what proportion of individuals in the population each receives good reputations, defined as goodness here) becomes complex, and thus has been studied mainly with numerical simulations. The present study gives a theoretical analysis of such complex reputation structure. We formulate the time change of goodness of individuals caused by updates of reputations among individuals. By considering a large population, we derive dynamics of the frequency distribution of goodnesses. An equilibrium state of the dynamics is approximated by a summation of Gaussian functions. We demonstrate that the theoretical solution well fits the numerical calculation. From the theoretical solution, we obtain a new interpretation of the complex reputation structure. This study provides a novel mathematical basis for cutting-edge studies on indirect reciprocity.

physics.soc-ph

Exploitation by asymmetry of information reference in coevolutionary learning in prisoner's dilemma game

Mutual relationships, such as cooperation and exploitation, are the basis of human and other biological societies. The foundations of these relationships are rooted in the decision making of individuals, and whether they choose to be selfish or altruistic. How individuals choose their behaviors can be analyzed using a strategy optimization process in the framework of game theory. Previous studies have shown that reference to individuals' previous actions plays an important role in their choice of strategies and establishment of social relationships. A fundamental question remains as to whether an individual with more information can exploit another who has less information when learning the choice of strategies. Here we demonstrate that a player using a memory-one strategy, who can refer to their own previous action and that of their opponent, can be exploited by a reactive player, who only has the information of the other player, based on mutual adaptive learning. This is counterintuitive because the former has more choice in strategies and can potentially obtain a higher payoff. We demonstrated this by formulating the learning process of strategy choices to optimize the payoffs in terms of coupled replicator dynamics and applying it to the prisoner's dilemma game. Further, we show that the player using a memory-one strategy, by referring to their previous experience, can sometimes act more generous toward the opponent's defection, thereby accepting the opponent's exploitation. Mainly, we found that through adaptive learning, a player with limited information usually exploits the player with more information, leading to asymmetric exploitation.

math.OC

Emergence of Exploitation as Symmetry Breaking in Iterated Prisoner's Dilemma

In society, mutual cooperation, defection, and asymmetric exploitative relationships are common. Whereas cooperation and defection are studied extensively in the literature on game theory, asymmetric exploitative relationships between players are little explored. In a recent study, Press and Dyson demonstrate that if only one player can learn about the other, asymmetric exploitation is achieved in the prisoner's dilemma game. In contrast, however, it is unknown whether such one-way exploitation is stably established when both players learn about each other symmetrically and try to optimize their payoffs. Here, we first formulate a dynamical system that describes the change in a player's probabilistic strategy with reinforcement learning to obtain greater payoffs, based on the recognition of the other player. By applying this formulation to the standard prisoner's dilemma game, we numerically and analytically demonstrate that an exploitative relationship can be achieved despite symmetric strategy dynamics and symmetric rule of games. This exploitative relationship is stable, even though the exploited player, who receives a lower payoff than the exploiting player, has optimized the own strategy. Whether the final equilibrium state is mutual cooperation, defection, or exploitation, crucially depends on the initial conditions: Punishment against a defector oscillates between the players, and thus a complicated basin structure to the final equilibrium appears. In other words, slight differences in the initial state may lead to drastic changes in the final state. Considering the generality of the result, this study provides a new perspective on the origin of exploitation in society.

math.OC

Functional Dynamics by Intention Recognition in Iterated Games

Intention recognition is an important characteristic of intelligent agents. In their interactions with others, they try to read others' intentions and make an image of others to choose their actions accordingly. While the way in which players choose their actions depending on such intentions has been investigated in game theory, how dynamic changes in intentions by mutually reading others' intentions are incorporated into game theory has not been explored. We present a novel formulation of game theory in which players read others' intentions and change their own through an iterated game. Here, intention is given as a function of the other's action and the own action to be taken accordingly as the dependent variable, while the mutual recognition of intention is represented as the functional dynamics. It is shown that a player suffers no disadvantage when he/she recognizes the other's intention, whereas the functional dynamics reach equilibria in which both players' intentions are optimized. These cover a classical Nash and Stackelberg equilibria but we extend them in this study: Novel equilibria exist depending on the degree of mutual recognition. Moreover, the degree to which each player recognizes the other can also differ. This formulation is applied to resource competition, duopoly, and prisoner's dilemma games. For example, in the resource competition game with player-dependent capacity on gaining the resource, the superior player's recognition leads to the exploitation of the other, while the inferior player's recognition leads to cooperation through which both players' payoffs increase.

math.OC

Hierarchical Prisoner's Dilemma in Hierarchical Public-Goods Game

The dilemma in cooperation is one of the major concerns in game theory. In a public-goods game, each individual pays a cost for cooperation, or to prevent defection, and receives a reward from the collected cost in a group. Thus, defection is beneficial for each individual, while cooperation is beneficial for the group. Now, groups (say, countries) consisting of individual players also play games. To study such a multi-level game, we introduce a hierarchical public-goods (HPG) game in which two groups compete for finite resources by utilizing costs collected from individuals in each group. Analyzing this HPG game, we found a hierarchical prisoner's dilemma, in which groups choose the defection policy (say, armaments) as a Nash strategy to optimize each group's benefit, while cooperation optimizes the total benefit. On the other hand, for each individual within a group, refusing to pay the cost (say, tax) is a Nash strategy, which turns to be a cooperation policy for the group, thus leading to a hierarchical dilemma. Here, the reward received by one group increases with the population, as does the collected cost. In spite of this, we find that there exists an optimal group size that maximizes its payoff. Furthermore, when the population asymmetry between two groups is large, a smaller group will choose a cooperation policy (say, disarmament) to avoid excessive response from the larger group, which leads to the resolution of the prisoner's dilemma between the groups. The relevance of the HPG game to policy selection in society and the optimal size in human or animal groups are discussed accordingly.

physics.soc-ph