SearcharxivSearch

arXiv subjects

Jiqiang Zhang

Publications and source records attributed to Jiqiang Zhang.

At least 19 recordsLinked to original sources

Evolution of cooperation with Q-learning: how much information do we need?

Cooperation is ubiquitous in both natural and human societies, yet its evolutionary basis remains a major challenge. A long-standing puzzle is whether having more information leads to better decision-making and thus a higher level of cooperation. To address this question, we adopt a recently developed reinforcement learning framework in which individuals learn through trial and error to maximize cumulative rewards - a paradigm that has successfully explained diverse emergent patterns in human behaviors. Specifically, we equip a structured population with the Q-learning algorithm and systematically vary the size of the interactive neighborhood, which serves as a proxy for perceived information. Interestingly, we observe a non-monotonic relationship between cooperation prevalence and neighborhood size in both two-dimensional square lattices and Barabasi-Albert scale-free networks. This inverted U-shaped dependence reveals that an optimal amount of information exists, yielding the highest level of cooperation. Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability. This balance allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload. Our findings challenge everyday intuition, suggesting that a proper amount of information - not more - is optimal for the emergence of cooperation.

physics.soc-ph

Emergence of cooperation: A reputation-modulated reinforcement learning

Reputation is widely recognized as a key mechanism for sustaining cooperation. However, most existing game-theoretic models treat reputation primarily as an external factor that modulates payoffs, interaction structures, or strategy update rules. In many social contexts, though, reputation operates primarily as information -- it shapes how individuals interpret their own experiences and assess the behavior of others. To bridge this gap, we propose a spatial prisoner's dilemma game grounded in the reinforcement learning paradigm, in which agents equipped with Q-learning integrate both individual and social information via a locally defined reputation metric to guide their decisions. Our results reveal that reputation-modulated learning significantly promotes the emergence of cooperative behavior, and we observe a discontinuous phase transition from full cooperation to full defection as the temptation increases. Cooperation spreads through the nucleation of cooperative clusters, whereas the disintegration of these clusters drives the system into an absorbing state of complete defection. Overall, this study demonstrates that reputation facilitates cooperation not only by providing direct incentives but also by reshaping the social information landscape that agents rely on for learning and adaptation.

physics.soc-ph

Decoding species coexistence: A reinforcement learning perspective

A central goal in ecology is to understand how biodiversity is maintained. Previous theoretical works have employed the rock-paper-scissors (RPS) game as a toy model, demonstrating that population mobility is crucial in determining the species' coexistence. One key prediction is that biodiversity is jeopardized and eventually lost when mobility exceeds a certain value--a conclusion at odds with empirical observations of highly mobile species coexisting in nature. To address this discrepancy, we introduce a reinforcement learning framework and study a spatial RPS model, where individual mobility is adaptively regulated via a Q-learning algorithm rather than held fixed. Our results show that all three species can coexist stably, with extinction probabilities remaining low across a broad range of baseline migration rates. Mechanistic analysis reveals that individuals develop two behavioral tendencies: survival priority (escaping from predators) and predation priority (remaining near prey). While species coexistence emerges from the balance of the two tendencies, their imbalance jeopardizes biodiversity. Notably, there is a symmetry-breaking of action preference in a particular state that is responsible for the divergent species densities. Furthermore, when Q-learning species interact with fixed-mobility counterparts, those with adaptive mobility exhibit a significant evolutionary advantage. Our study suggests that reinforcement learning may offer a promising new perspective for uncovering the mechanisms of biodiversity and informing conservation strategies.

q-bio.PE

A brief review of evolutionary game dynamics in the reinforcement learning paradigm

Cooperation, fairness, trust, and resource coordination are cornerstones of modern civilization, yet their emergence remains inadequately explained by the persistent discrepancies between theoretical predictions and behavioral experiments. Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models, which assumes individuals merely copy successful neighbors according to predetermined, fixed rules. This review examines recent advances in evolutionary game dynamics that employ reinforcement learning (RL) as an alternative paradigm. In RL, individuals learn through trial and error and introspectively refine their strategies based on environmental feedback. We begin by introducing key concepts in evolutionary game theory and the two learning paradigms, then synthesize progress in applying RL to elucidate cooperation, trust, fairness, optimal resource coordination, and ecological dynamics. Collectively, these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.

q-bio.PE

Optimal coordination of resources: A solution from reinforcement learning

Efficient allocation is important in nature and human society, where individuals frequently compete for limited resources. The Minority Game (MG) is perhaps the simplest toy model to address this issue. However, most previous solutions assume that the strategies are provided a priori and static, failing to capture their adaptive nature. Here, we introduce the reinforcement learning (RL) paradigm to MG, where individuals adjust decisions based on accumulated experience and expected rewards dynamically. We find that this RL framework achieves optimal resource coordination when individuals balance the exploitation of experience with random exploration. Yet, the imbalanced strategies of the two lead to suboptimal partial coordination or even anti-coordination. Our mechanistic analysis reveals a symmetry-breaking in action preferences at the optimum, offering a fresh solution to the MG and new insights into the resource allocation problem.

physics.soc-ph

Evolution of cooperation with Q-learning: the impact of information perception

The inherent complexity of human beings manifests in a remarkable diversity of responses to intricate environments, enabling us to approach problems from varied perspectives. However, in the study of cooperation, existing research within the reinforcement learning framework often assumes that individuals have access to identical information when making decisions, which contrasts with the reality that individuals frequently perceive information differently. In this study, we employ the Q-learning algorithm to explore the impact of information perception on the evolution of cooperation in a two-person Prisoner's Dilemma game. We demonstrate that the evolutionary processes differ significantly across three distinct information perception scenarios, highlighting the critical role of information structure in the emergence of cooperation. Notably, the asymmetric information scenario reveals a complex dynamical process, including the emergence, breakdown, and reconstruction of cooperation, mirroring psychological shifts observed in human behavior. Our findings underscore the importance of information structure in fostering cooperation, offering new insights into the establishment of stable cooperative relationships among humans.

q-bio.PE

Evolution of cooperation in a bimodal mixture of conditional cooperators

Extensive behavioral experiments reveal that conditional cooperation is a prevalent phenomenon. Previous game-theoretical studies have predominantly relied on hard-manner models, where cooperation is triggered only upon reaching a specific threshold. However, this approach contrasts with the observed flexibility of human behaviors, where individuals adapt their strategies dynamically based on their surroundings. To capture this adaptability, we introduce a soft form of conditional cooperation by integrating the Q-learning algorithm from reinforcement learning. In this form, players not only reciprocate mutual cooperation but may also defect in highly cooperative environments or cooperate in less cooperative settings to maximize rewards. To explore the effects of hard and soft conditional cooperators, we examine their interactions in two scenarios: structural mixture (SM) and probabilistic mixture (PM), where the two behavioral modes are fixed and probabilistically adopted, respectively. In SM, hard conditional cooperators enhance cooperation when the threshold is low but hinder it otherwise. Surprisingly, in PM, the cooperation prevalence exhibits two first-order phase transitions as the probability is varied, leading to high, low, and vanishing levels of cooperation. Analysis of Q-tables offers insights into the "psychological shifts" of soft conditional cooperators and the overall evolutionary dynamics. Model extensions confirm the robustness of our findings. These results highlight the novel complexities arising from the diversity of conditional cooperators.

q-bio.PE

Decoding fairness: a reinforcement learning perspective

Behavioral experiments on the ultimatum game (UG) reveal that we humans prefer fair acts, which contradicts the prediction made in orthodox Economics. Existing explanations, however, are mostly attributed to exogenous factors within the imitation learning framework. Here, we adopt the reinforcement learning paradigm, where individuals make their moves aiming to maximize their accumulated rewards. Specifically, we apply Q-learning to UG, where each player is assigned two Q-tables to guide decisions for the roles of proposer and responder. In a two-player scenario, fairness emerges prominently when both experiences and future rewards are appreciated. In particular, the probability of successful deals increases with higher offers, which aligns with observations in behavioral experiments. Our mechanism analysis reveals that the system undergoes two phases, eventually stabilizing into fair or rational strategies. These results are robust when the rotating role assignment is replaced by a random or fixed manner, or the scenario is extended to a latticed population. Our findings thus conclude that the endogenous factor is sufficient to explain the emergence of fairness, exogenous factors are not needed.

cs.LG

Mitigating Exposure Bias in Score-Based Generation of Molecular Conformations

Molecular conformation generation poses a significant challenge in the field of computational chemistry. Recently, Diffusion Probabilistic Models (DPMs) and Score-Based Generative Models (SGMs) are effectively used due to their capacity for generating accurate conformations far beyond conventional physics-based approaches. However, the discrepancy between training and inference rises a critical problem known as the exposure bias. While this issue has been extensively investigated in DPMs, the existence of exposure bias in SGMs and its effective measurement remain unsolved, which hinders the use of compensation methods for SGMs, including ConfGF and Torsional Diffusion as the representatives. In this work, we first propose a method for measuring exposure bias in SGMs used for molecular conformation generation, which confirms the significant existence of exposure bias in these models and measures its value. We design a new compensation algorithm Input Perturbation (IP), which is adapted from a method originally designed for DPMs only. Experimental results show that by introducing IP, SGM-based molecular conformation models can significantly improve both the accuracy and diversity of the generated conformations. Especially by using the IP-enhanced Torsional Diffusion model, we achieve new state-of-the-art performance on the GEOM-Drugs dataset and are on par on GEOM-QM9. We provide the code publicly at https://github.com/jia-975/torsionalDiff-ip.

cs.LG

Multiplex Graph Contrastive Learning with Soft Negatives

Graph Contrastive Learning (GCL) seeks to learn nodal or graph representations that contain maximal consistent information from graph-structured data. While node-level contrasting modes are dominating, some efforts commence to explore consistency across different scales. Yet, they tend to lose consistent information and be contaminated by disturbing features. Here, we introduce MUX-GCL, a novel cross-scale contrastive learning paradigm that utilizes multiplex representations as effective patches. While this learning mode minimizes contaminating noises, a commensurate contrasting strategy using positional affinities further avoids information loss by correcting false negative pairs across scales. Extensive downstream experiments demonstrate that MUX-GCL yields multiple state-of-the-art results on public datasets. Our theoretical analysis further guarantees the new objective function as a stricter lower bound of mutual information of raw input features and output embeddings, which rationalizes this paradigm. Code is available at https://github.com/MUX-GCL/Code.

cs.LG

Evolution of cooperation in the public goods game with Q-learning

Recent paradigm shifts from imitation learning to reinforcement learning (RL) is shown to be productive in understanding human behaviors. In the RL paradigm, individuals search for optimal strategies through interaction with the environment to make decisions. This implies that gathering, processing, and utilizing information from their surroundings are crucial. However, existing studies typically study pairwise games such as the prisoners' dilemma and employ a self-regarding setup, where individuals play against one opponent based solely on their own strategies, neglecting the environmental information. In this work, we investigate the evolution of cooperation with the multiplayer game -- the public goods game using the Q-learning algorithm by leveraging the environmental information. Specifically, the decision-making of players is based upon the cooperation information in their neighborhood. Our results show that cooperation is more likely to emerge compared to the case of imitation learning by using Fermi rule. Of particular interest is the observation of an anomalous non-monotonic dependence which is revealed when voluntary participation is further introduced. The analysis of the Q-table explains the mechanisms behind the cooperation evolution. Our findings indicate the fundamental role of environment information in the RL paradigm to understand the evolution of cooperation, and human behaviors in general.

q-bio.PE

Catalytic evolution of cooperation in a population with behavioural bimodality

The remarkable adaptability of humans in response to complex environments is often demonstrated by the context-dependent adoption of different behavioral modes. However, the existing game-theoretic studies mostly focus on the single-mode assumption, and the impact of this behavioral multimodality on the evolution of cooperation remains largely unknown. Here, we study how cooperation evolves in a population with two behavioral modes. Specifically, we incorporate Q-learning and Tit-for-Tat (TFT) rules into our toy model, where prisoner's dilemma game is played and we investigate the impact of the mode mixture on the evolution of cooperation. While players in Q-learning mode aim to maximize their accumulated payoffs, players within TFT mode repeat what their neighbors have done to them. In a structured mixing implementation where the updating rule is fixed for each individual, we find that the mode mixture greatly promotes the overall cooperation prevalence. The promotion is even more significant in the probabilistic mixing, where players randomly select one of the two rules at each step. Finally, this promotion is robust when players are allowed to adaptively choose the two modes by real-time comparison. In all three scenarios, players within the Q-learning mode act as catalyzer that turns the TFT players to be more cooperative, and as a result drive the whole population to be highly cooperative. The analysis of Q-tables explains the underlying mechanism of cooperation promotion, which captures the ``psychologic evolution" in the players' mind. Our study indicates that the variety of behavioral modes is non-negligible, and could be crucial to clarify the emergence of cooperation in the real world.

q-bio.PE

GraphPub: Generation of Differential Privacy Graph with High Availability

In recent years, with the rapid development of graph neural networks (GNN), more and more graph datasets have been published for GNN tasks. However, when an upstream data owner publishes graph data, there are often many privacy concerns, because many real-world graph data contain sensitive information like person's friend list. Differential privacy (DP) is a common method to protect privacy, but due to the complex topological structure of graph data, applying DP on graphs often affects the message passing and aggregation of GNN models, leading to a decrease in model accuracy. In this paper, we propose a novel graph edge protection framework, graph publisher (GraphPub), which can protect graph topology while ensuring that the availability of data is basically unchanged. Through reverse learning and the encoder-decoder mechanism, we search for some false edges that do not have a large negative impact on the aggregation of node features, and use them to replace some real edges. The modified graph will be published, which is difficult to distinguish between real and false data. Sufficient experiments prove that our framework achieves model accuracy close to the original graph with an extremely low privacy budget.

cs.SI

Emergence of cooperation under punishment: A reinforcement learning perspective

Punishment is a common tactic to sustain cooperation and has been extensively studied for a long time. While most of previous game-theoretic work adopt the imitation learning where players imitate the strategies who are better off, the learning logic in the real world is often much more complex. In this work, we turn to the reinforcement learning paradigm, where individuals make their decisions based upon their past experience and long-term returns. Specifically, we investigate the Prisoners' dilemma game with Q-learning algorithm, and cooperators probabilistically pose punishment on defectors in their neighborhood. Interestingly, we find that punishment could lead to either continuous or discontinuous cooperation phase transitions, and the nucleation process of cooperation clusters is reminiscent of the liquid-gas transition. The uncovered first-order phase transition indicates that great care needs to be taken when implementing the punishment compared to the continuous scenario.

q-bio.PE

Decoding trust: A reinforcement learning perspective

Behavioral experiments on the trust game have shown that trust and trustworthiness are universal among human beings, contradicting the prediction by assuming \emph{Homo economicus} in orthodox Economics. This means some mechanism must be at work that favors their emergence. Most previous explanations however need to resort to some factors based upon imitative learning, a simple version of social learning. Here, we turn to the paradigm of reinforcement learning, where individuals update their strategies by evaluating the long-term return through accumulated experience. Specifically, we investigate the trust game with the Q-learning algorithm, where each participant is associated with two evolving Q-tables that guide one's decision making as trustor and trustee respectively. In the pairwise scenario, we reveal that high levels of trust and trustworthiness emerge when individuals appreciate both their historical experience and returns in the future. Mechanistically, the evolution of the Q-tables shows a crossover that resembles human's psychological changes. We also provide the phase diagram for the game parameters, where the boundary analysis is conducted. These findings are robust when the scenario is extended to a latticed population. Our results thus provide a natural explanation for the emergence of trust and trustworthiness without external factors involved. More importantly, the proposed paradigm shows the potential in deciphering many puzzles in human behaviors.

q-bio.PE

A Generative Hypergraph Model for Double Heterogeneity

While network science has become an indispensable tool for studying complex systems, the conventional use of pairwise links often shows limitations in describing high-order interactions properly. Hypergraphs, where each edge can connect more than two nodes, have thus become a new paradigm in network science. Yet, we are still in lack of models linking network growth and hyperedge expansion, both of which are commonly observable in the real world. Here, we propose a generative hypergraph model by employing the preferential attachment mechanism in both nodes and hyperedge formation. The model can produce bi-heterogeneity, exhibiting scale-free distributions in both hyperdegree and hyperedge size. We provide a mean-field treatment that gives the expression of the two scaling exponents, which agree with the numerical simulations. Our model may help to understand the networked systems showing both types of heterogeneity and facilitate the study of complex dynamics thereon.

physics.soc-ph

Emergence of cooperation in a population with bimodal response behaviors

We human beings show remarkable adaptability in response to complex surroundings, we adopt different behavioral modes at different occasions, such response multimodality is critical to our survival. Yet, how this behavioral multimodality affects the evolution of cooperation remains largely unknown. Here we build a toy model to address this issue by considering a population with bimodal response behaviors, or specifically, with the Fermi and Tit-for-tat updating rules. While the former rule tends to imitate the strategies of those neighbors who are doing well, the latter repeats what their neighbors did to them. In a structural mixing implementation, where the updating rule is fixed for each individual, we find that a moderate mode mixture unexpectedly boosts the overall cooperation level of the population. The boost is even more pronounced in the probabilistic mixing, where each individual randomly chooses one of the two modes at each step, and full cooperation is seen in a wide range. These findings are robust to the underlying topology of the population. Our mean-field treatment reveals that the cooperation prevalence within the players with the Fermi rule linearly increases with the fraction of TFT players and explains the non-monotonic dependence in the structural mixing. Our study shows that the diversity in response behaviors may help to explain the emergence of cooperation in realistic contexts.

physics.soc-ph

Pinning control of social fairness in the Ultimatum game

Decent social fairness is highly desired both for socio-economic activities and individuals, as it is one of the cornerstones of our social welfare and sustainability. How to effectively promote the level of fairness thus becomes a significant issue to be addressed. Here, by adopting a pinning control procedure, we find that when a very small fraction of individuals are pinned to be fair players in the Ultimatum Game, the whole population unexpectedly evolves into the full fairness level. The basic observations are quite robust in homogeneous networks, but the converging time as a function of the pinning number shows different laws for different underlying topologies. For heterogeneous networks, this leverage effect is even more pronounced that one hub node is sufficient for the aim, and a periodic on-off control procedure can be applied to further save the control cost. Intermittent failures are seen when the pinning control is marginally strong, our statistical analysis indicates some sort of criticality. Our work suggests that the pinning control procedure could potentially be a good strategy to promote the social fairness for some real scenarios when necessary.

physics.soc-ph