SearcharxivSearch

arXiv subjects

Shengfeng Deng

Publications and source records attributed to Shengfeng Deng.

At least 19 recordsLinked to original sources

Evolution of cooperation with Q-learning: how much information do we need?

Cooperation is ubiquitous in both natural and human societies, yet its evolutionary basis remains a major challenge. A long-standing puzzle is whether having more information leads to better decision-making and thus a higher level of cooperation. To address this question, we adopt a recently developed reinforcement learning framework in which individuals learn through trial and error to maximize cumulative rewards - a paradigm that has successfully explained diverse emergent patterns in human behaviors. Specifically, we equip a structured population with the Q-learning algorithm and systematically vary the size of the interactive neighborhood, which serves as a proxy for perceived information. Interestingly, we observe a non-monotonic relationship between cooperation prevalence and neighborhood size in both two-dimensional square lattices and Barabasi-Albert scale-free networks. This inverted U-shaped dependence reveals that an optimal amount of information exists, yielding the highest level of cooperation. Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability. This balance allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload. Our findings challenge everyday intuition, suggesting that a proper amount of information - not more - is optimal for the emergence of cooperation.

physics.soc-ph

A brief review of evolutionary game dynamics in the reinforcement learning paradigm

Cooperation, fairness, trust, and resource coordination are cornerstones of modern civilization, yet their emergence remains inadequately explained by the persistent discrepancies between theoretical predictions and behavioral experiments. Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models, which assumes individuals merely copy successful neighbors according to predetermined, fixed rules. This review examines recent advances in evolutionary game dynamics that employ reinforcement learning (RL) as an alternative paradigm. In RL, individuals learn through trial and error and introspectively refine their strategies based on environmental feedback. We begin by introducing key concepts in evolutionary game theory and the two learning paradigms, then synthesize progress in applying RL to elucidate cooperation, trust, fairness, optimal resource coordination, and ecological dynamics. Collectively, these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.

q-bio.PE

Supervised and unsupervised learning with numerical computation for the Wolfram cellular automata

The local rules of Wolfram cellular automata with one-dimensional three-cell neighborhoods are represented by eight-bit binary that encode deterministic update rules. These automata are widely utilized to investigate self-organization phenomena and the dynamics of complex systems. In this work, we employ numerical simulations and computational methods to investigate the asymptotic density and dynamical evolution mechanisms in Wolfram automata. We apply both supervised and unsupervised learning methods to identify the configurations associated with different Wolfram rules. Furthermore, we explore alternative initial conditions under which certain Wolfram rules generate similar fractal patterns over time, even when starting from a single active site. Our results reveal the relationship between the asymptotic density and the initial density of selected rules. The supervised learning methods effectively identify the configurations of various Wolfram rules, while unsupervised methods like principal component analysis and autoencoders can approximately cluster configurations of different Wolfram rules into distinct groups, yielding results that align well with simulated density outputs.

physics.comp-ph

Decoding species coexistence: A reinforcement learning perspective

A central goal in ecology is to understand how biodiversity is maintained. Previous theoretical works have employed the rock-paper-scissors (RPS) game as a toy model, demonstrating that population mobility is crucial in determining the species' coexistence. One key prediction is that biodiversity is jeopardized and eventually lost when mobility exceeds a certain value--a conclusion at odds with empirical observations of highly mobile species coexisting in nature. To address this discrepancy, we introduce a reinforcement learning framework and study a spatial RPS model, where individual mobility is adaptively regulated via a Q-learning algorithm rather than held fixed. Our results show that all three species can coexist stably, with extinction probabilities remaining low across a broad range of baseline migration rates. Mechanistic analysis reveals that individuals develop two behavioral tendencies: survival priority (escaping from predators) and predation priority (remaining near prey). While species coexistence emerges from the balance of the two tendencies, their imbalance jeopardizes biodiversity. Notably, there is a symmetry-breaking of action preference in a particular state that is responsible for the divergent species densities. Furthermore, when Q-learning species interact with fixed-mobility counterparts, those with adaptive mobility exhibit a significant evolutionary advantage. Our study suggests that reinforcement learning may offer a promising new perspective for uncovering the mechanisms of biodiversity and informing conservation strategies.

q-bio.PE

Decoding fairness: a reinforcement learning perspective

Behavioral experiments on the ultimatum game (UG) reveal that we humans prefer fair acts, which contradicts the prediction made in orthodox Economics. Existing explanations, however, are mostly attributed to exogenous factors within the imitation learning framework. Here, we adopt the reinforcement learning paradigm, where individuals make their moves aiming to maximize their accumulated rewards. Specifically, we apply Q-learning to UG, where each player is assigned two Q-tables to guide decisions for the roles of proposer and responder. In a two-player scenario, fairness emerges prominently when both experiences and future rewards are appreciated. In particular, the probability of successful deals increases with higher offers, which aligns with observations in behavioral experiments. Our mechanism analysis reveals that the system undergoes two phases, eventually stabilizing into fair or rational strategies. These results are robust when the rotating role assignment is replaced by a random or fixed manner, or the scenario is extended to a latticed population. Our findings thus conclude that the endogenous factor is sufficient to explain the emergence of fairness, exogenous factors are not needed.

cs.LG

Evolution of cooperation in the public goods game with Q-learning

Recent paradigm shifts from imitation learning to reinforcement learning (RL) is shown to be productive in understanding human behaviors. In the RL paradigm, individuals search for optimal strategies through interaction with the environment to make decisions. This implies that gathering, processing, and utilizing information from their surroundings are crucial. However, existing studies typically study pairwise games such as the prisoners' dilemma and employ a self-regarding setup, where individuals play against one opponent based solely on their own strategies, neglecting the environmental information. In this work, we investigate the evolution of cooperation with the multiplayer game -- the public goods game using the Q-learning algorithm by leveraging the environmental information. Specifically, the decision-making of players is based upon the cooperation information in their neighborhood. Our results show that cooperation is more likely to emerge compared to the case of imitation learning by using Fermi rule. Of particular interest is the observation of an anomalous non-monotonic dependence which is revealed when voluntary participation is further introduced. The analysis of the Q-table explains the mechanisms behind the cooperation evolution. Our findings indicate the fundamental role of environment information in the RL paradigm to understand the evolution of cooperation, and human behaviors in general.

q-bio.PE

Evolution of cooperation with Q-learning: the impact of information perception

The inherent complexity of human beings manifests in a remarkable diversity of responses to intricate environments, enabling us to approach problems from varied perspectives. However, in the study of cooperation, existing research within the reinforcement learning framework often assumes that individuals have access to identical information when making decisions, which contrasts with the reality that individuals frequently perceive information differently. In this study, we employ the Q-learning algorithm to explore the impact of information perception on the evolution of cooperation in a two-person Prisoner's Dilemma game. We demonstrate that the evolutionary processes differ significantly across three distinct information perception scenarios, highlighting the critical role of information structure in the emergence of cooperation. Notably, the asymmetric information scenario reveals a complex dynamical process, including the emergence, breakdown, and reconstruction of cooperation, mirroring psychological shifts observed in human behavior. Our findings underscore the importance of information structure in fostering cooperation, offering new insights into the establishment of stable cooperative relationships among humans.

q-bio.PE

Supervised and unsupervised learning of directed percolation

Machine learning (ML) has been well applied to studying equilibrium phase transition models, by accurately predicating critical thresholds and some critical exponents. Difficulty will be raised, however, for integrating ML into non-equilibrium phase transitions. The extra dimension in a given non-equilibrium system, namely time, can greatly slow down the procedure towards the steady state. In this paper we find that by using some simple techniques of ML, non-steady state configurations of directed percolation (DP) suffice to capture its essential critical behaviors in both (1+1) and (2+1) dimensions. With the supervised learning method, the framework of our binary classification neural networks can identify the phase transition threshold, as well as the spatial and temporal correlation exponents. The characteristic time $t_{c}$, specifying the transition from active phases to absorbing ones, is also a major product of the learning. Moreover, we employ the convolutional autoencoder, an unsupervised learning technique, to extract dimensionality reduction representations and cluster configurations of (1+1) bond DP. It is quite appealing that such a method can yield a reasonable estimation of the critical point.

cond-mat.stat-mech

Machine learning of pair-contact process with diffusion

The pair-contact process with diffusion (PCPD), a generalized model of the ordinary pair-contact process (PCP) without diffusion, exhibits a continuous absorbing phase transition. Unlike the PCP, whose nature of phase transition is clearly classified into the directed percolation (DP) universality class, the model of PCPD has been controversially discussed since its infancy. To our best knowledge, there is so far no consensus on whether the phase transition of the PCPD falls into the unknown university classes or else conveys a new kind of non-equilibrium phase transition. In this paper, both unsupervised and supervised learning are employed to study the PCPD with scrutiny. Firstly, two unsupervised learning methods, principal component analysis (PCA) and autoencoder, are taken. Our results show that both methods can cluster the original configurations of the model and provide reasonable estimates of thresholds. Therefore, no matter whether the non-equilibrium lattice model is a random process of unitary (for instance the DP) or binary (for instance the PCP), or whether it contains the diffusion motion of particles, unsupervised leaning can capture the essential, hidden information. Beyond that, supervised learning is also applied to learning the PCPD at different diffusion rates. We proposed a more accurate numerical method to determine the spatial correlation exponent $ν_{\perp}$, which, to a large degree, avoids the uncertainty of data collapses through naked eyes. Our extensive calculations reveal that $ν_{\perp}$ of PCPD depends continuously on the diffusion rate $D$, which supports the viewpoint that the PCPD may lead to a new type of absorbing phase transition.

cond-mat.stat-mech

Supervised, semi-supervised, and unsupervised learning of the Domany-Kinzel model

The Domany Kinzel (DK) model encompasses several types of non-equilibrium phase transitions, depending on the selected parameters. We apply supervised, semi-supervised, and unsupervised learning methods to studying the phase transitions and critical behaviors of the (1 + 1)-dimensional DK model. The supervised and the semi-supervised learning methods permit the estimations of the critical points, the spatial and temporal correlation exponents, concerning labelled and unlabelled DK configurations, respectively. Furthermore, we also predict the critical points by employing principal component analysis (PCA) and autoencoder. The PCA and autoencoder can produce results in good agreement with simulated particle number density.

physics.comp-ph

Dynamical heterogeneity and universality of power-grids

While weak, tuned asymmetry can improve, strong heterogeneity destroys synchronization in the electric power system. We study the level of heterogeneity, by comparing large high voltage (HV) power-grids of Europe and North America. We provide an analysis of power capacities and loads of various energy sources from the databases and found heavy tailed distributions with similar characteristics. Graph topological measures, community structures also exhibit strong similarities, while the cable admittance distributions can be well fitted with the same power-laws (PL), related to the length distributions. The community detection analysis shows the level of synchronization in different domains of the European HV power grids, by solving a set of swing equations. We provide numerical evidence for frustrated synchronization and Chimera states and point out the relation of topology and level of synchronization in the subsystems. We also provide empirical data analysis of the frequency heterogeneities within the Hungarian HV network and find q-Gaussian distributions related to super-statistics of time-lagged fluctuations, which agree well with former results on the Nordic Grid.

physics.soc-ph

Revisiting and modeling power-law distributions in empirical outage data of power systems

The size distribution of planned and forced outages and following restoration times in power systems have been studied for almost two decades and has drawn great interest as they display heavy tails. Understanding of this phenomenon has been done by various threshold models, which are self-tuned at their critical points, but as many papers pointed out, explanations are intuitive, and more empirical data is needed to support hypotheses. In this paper, the authors analyze outage data collected from various public sources to calculate the outage energy and outage duration exponents of possible power-law fits. Temporal thresholds are applied to identify crossovers from initial short-time behavior to power-law tails. We revisit and add to the possible explanations of the uniformness of these exponents. By performing power spectral analyses on the outage event time series and the outage duration time series, it is found that, on the one hand, while being overwhelmed by white noise, outage events show traits of self-organized criticality (SOC), which may be modeled by a crossover from random percolation to directed percolation branching process with dissipation, coupled to a conserved density. On the other hand, in responses to outages, the heavy tails in outage duration distributions could be a consequence of the highly optimized tolerance (HOT) mechanism, based on the optimized allocation of maintenance resources.

cond-mat.stat-mech

Chimera states in neural networks and power systems

Partial, frustrated synchronization and chimera-like states are expected to occur in Kuramoto-like models if the spectral dimension of the underlying graph is low: $d_s < 4$. We provide numerical evidence that this really happens in case of the high-voltage power grid of Europe ($d_s < 2$), a large human connectome (KKI113) and in case of the largest, exactly known brain network corresponding to the fruit-fly (FF) connectome ($d_s < 4$), even though their graph dimensions are much higher, i.e.: $d^{EU}_g\simeq 2.6(1)$ and $d^{FF}_g\simeq 5.4(1)$, $d^{\mathrm{KKI113}}_g\simeq 3.4(1)$. We provide local synchronization results of the first- and second-order (Shinomoto) Kuramoto models by numerical solutions on the FF and the European power-grid graphs, respectively, and show the emergence of \red{chimera-like} patterns on the graph community level as well as by the local order parameters.

cond-mat.stat-mech

Synchronization transitions on connectome graphs with external force

We investigate the synchronization transition of the Shinomoto-Kuramoto model on networks of the fruit-fly and two large human connectomes. This model contains a force term, thus is capable of describing critical behavior in the presence of external excitation. By numerical solution we determine the crackling noise durations with and without thermal noise and show extended non-universal scaling tails characterized by $2< τ_t < 2.8$, in contrast with the Hopf transition of the Kuramoto model, without the force $τ_t=3.1(1)$. Comparing the phase and frequency order parameters we find different transition points and fluctuations peaks as in case of the Kuramoto model. Using the local order parameter values we also determine the Hurst (phase) and $β$ (frequency) exponents and compare them with recent experimental results obtained by fMRI. We show that these exponents, characterizing the auto-correlations are smaller in the excited system than in the resting state and exhibit module dependence.

cond-mat.dis-nn

Critical behavior of the diffusive susceptible-infected-recovered model

The critical behavior of the non-diffusive susceptible-infected-recovered model on lattices had been well established in virtue of its duality symmetry. By performing simulations and scaling analyses for the diffusive variant on the two-dimensional lattice, we show that diffusion for all agents, while rendering this symmetry destroyed, constitutes a singular perturbation that induces asymptotically distinct dynamical and stationary critical behavior from the non-diffusive model. In particular, the manifested crossover behavior in the effective mean-square radius exponents reveals that slow crossover behavior in general diffusive multi-species reaction systems may be ascribed to the interference of multiple length scales and timescales at early times.

cond-mat.stat-mech

Synchronization transition of the second-order Kuramoto model on lattices

The second-order Kuramoto equation describes synchronization of coupled oscillators with inertia, which occur in power grids for example. Contrary to the first-order Kuramoto equation it's synchronization transition behavior is much less known. In case of Gaussian self-frequencies it is discontinuous, in contrast to the continuous transition for the first-order Kuramoto equation. Here we investigate this transition on large 2d and 3d lattices and provide numerical evidence of hybrid phase transitions, that the oscillator phases $θ_i$, exhibit a crossover, while the frequency spread a real phase transition in 3d. Thus a lower critical dimension $d_l^O=2$ is expected for the frequencies and $d_l^R=4$ for the phases like in the massless case. We provide numerical estimates for the critical exponents, finding that the frequency spread decays as $\sim t^{-d/2}$ in case of aligned initial state of the phases in agreement with the linear approximation. However in 3d, in the case of initially random distribution of $θ_i$, we find a faster decay, characterized by $\sim t^{-1.8(1)}$ as the consequence of enhanced nonlinearities which appear by the random phase fluctuations.

cond-mat.stat-mech

Transfer learning of phase transitions in percolation and directed percolation

The latest advances of statistical physics have shown remarkable performance of machine learning in identifying phase transitions. In this paper, we apply domain adversarial neural network (DANN) based on transfer learning to studying non-equilibrium and equilibrium phase transition models, which are percolation model and directed percolation (DP) model, respectively. With the DANN, only a small fraction of input configurations (2d images) needs to be labeled, which is automatically chosen, in order to capture the critical point. To learn the DP model, the method is refined by an iterative procedure in determining the critical point, which is a prerequisite for the data collapse in calculating the critical exponent $ν_{\perp}$. We then apply the DANN to a two-dimensional site percolation with configurations filtered to include only the largest cluster which may contain the information related to the order parameter. The DANN learning of both models yields reliable results which are comparable to the ones from Monte Carlo simulations. Our study also shows that the DANN can achieve quite high accuracy at much lower cost, compared to the supervised learning.

cond-mat.stat-mech

Synchronization dynamics on the EU and US power grids

Dynamical simulation of the cascade failures on the EU and USA high-voltage power grids has been done via solving the second-order Kuramoto equation. We show that synchronization transition happens by increasing the global coupling parameter $K$ with metasatble states depending on the initial conditions so that hysteresis loops occur. We provide analytic results for the time dependence of frequency spread in the large $K$ approximation and by comparing it with numerics of $d=2,3$ lattices, we find agreement in the case of ordered initial conditions. However, different power-law (PL) tails occur, when the fluctuations are strong. After thermalizing the systems we allow a single line cut failure and follow the subsequent overloads with respect to threshold values $T$. The PDFs $p(N_f)$ of the cascade failures exhibit PL tails near the synchronization transition point $K_c$. Near $K_c$ the exponents of the PL-s for the US power grid vary with $T$ as $1.4 \le τ\le 2.1$, in agreement with the empirical blackout statistics, while on the EU power grid we find somewhat steeper PL-s characterized by $1.4 \le τ\le 2.4$. Below $K_c$ we find signatures of $T$-dependent PL-s, caused by frustrated synchronization, reminiscent of Griffiths effects. Here we also observe stability growth following the blackout cascades, similar to intentional islanding, but for $K > K_c$ this does not happen. For $T < T_c$, bumps appear in the PDFs with large mean values, known as "dragon king" blackout events. We also analyze the delaying/stabilizing effects of instantaneous feedback or increased dissipation and show how local synchronization behaves on geographic maps.

cond-mat.stat-mech