SearcharxivSearch

arXiv subjects

Christian Hilbe

Publications and source records attributed to Christian Hilbe.

At least 19 recordsLinked to original sources

LLMs struggle to simulate human belief updates in controlled environments

LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.

cs.CL

Characterisation of reactive Nash equilibria in repeated additive games

In this paper, we study reactive strategies in repeated additive games between two players with finitely many actions. Reactive strategies condition only on the opponent's previous action, making them one of the simplest ways players can respond to past interactions. Additive games include important models of cooperation, such as the donation game and games with a punishment option. We show that, for this class of games and strategies, the conditions for symmetric Nash equilibria reduce to a system of linear equalities and inequalities in the strategy parameters, allowing us to characterise all such equilibria. We establish a one-to-one correspondence between non-empty subsets S of the action set and equilibrium classes, which we call S-supporting equilibria. These are equilibria that use exactly the actions in S when playing against themselves. As a special case, we recover the well-known equalizer strategies as the equilibria supported on the entire action set. To assess which equilibrium classes are most evolutionarily relevant, we complement our analytical characterisation with simulations of social learning dynamics. We find that their prevalence is determined by two factors: how likely they are to be generated and how robust they are against invasion.

cs.GT

Adaptive dynamics of alternating Prisoner's Dilemma with memory N

The Prisoner's Dilemma is used as a model in processes involving reciprocity; however, its classical setup can be insufficient in settings where the symmetry of the simultaneous decision making is broken -- for example, in donor and recipient processes. In the alternating Prisoner's Dilemma model the two players take turns choosing their strategy. Assuming a finite memory setup, we establish the mathematical aspects of the adaptive dynamics of the alternating Prisoner's Dilemma, paying particular attention to the case of memory 1.

math.DS

Strategies of cooperation and defection in five large language models

Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leading models produce sensible strategies in the repeated prisoner's dilemma, which is the main metaphor of reciprocal cooperation. First, we measure the propensity of LLMs to cooperate in a neutral setting, without using language reminiscent of how this game is usually presented. We record to what extent LLMs implement Nash equilibria or other well-known strategy classes. Thereafter, we explore how LLMs adapt their strategies to changes in parameter values. We vary the game's continuation probability, the payoff values, and whether the total number of rounds is commonly known. We also study the effect of different framings. In each case, we test whether the adaptations of the LLMs are in line with basic intuition, theoretical predictions of evolutionary game theory, and experimental evidence from human participants. While all LLMs perform well in many of the tasks, none of them exhibit full consistency over all tasks. We also conduct tournaments between the inferred LLM strategies and study direct interaction between LLMs in games over ten rounds with a known or unknown last round. Our experiments shed light on how current LLMs instantiate reciprocal cooperation.

cs.CY

Exact conditions for evolutionary stability in indirect reciprocity under noise

Indirect reciprocity is a key mechanism for large-scale cooperation. This mechanism captures the insight that in part, people help others to build and maintain a good reputation. To enable such cooperation, appropriate social norms are essential. They specify how individuals should act based on each others' reputations, and how reputations are updated in response to individual actions. Although previous work has identified several norms that sustain cooperation, a complete analytical characterization of all evolutionarily stable norms remains lacking, especially when assessments or actions are noisy. In this study, we provide such a characterization for the public assessment regime. This characterization reproduces known results, such as the leading eight norms, but it extends to more general cases, allowing for various types of errors and additional actions including costly punishment. We also identify norms that impose a fixed payoff on any mutant strategy, analogous to the zero-determinant strategies in direct reciprocity. These results offer a rigorous foundation for understanding the evolution of cooperation through indirect reciprocity and the critical role of social norms.

q-bio.PE

Evolution of noisy learning in games

People make strategic decisions many times a day - during negotiations, when coordinating actions with others, or when choosing partners for cooperation. The resulting dynamics can be studied with learning theory and evolutionary game theory. These frameworks explore how people adapt their decisions over time, in light of how effective their strategies have been. The outcomes of such learning processes depend on how sensitive individuals are to the performance of their strategies. When they are more sensitive, they systematically favor strategies they deem more successful. When they are less sensitive, their learning process is noisier and more erratic. Traditionally, most models treat this sensitivity as a fixed parameter - like the "selection strength" parameter in evolutionary models. Instead, we study how strategies and sensitivities co-evolve. We find that the co-evolutionary endpoints depend on both the type of strategic interaction and the learning rule employed. In prisoner's dilemmas, we often observe sensitivities to increase indefinitely. But in snowdrift and stag-hunt games, sensitivities often converge to a finite value, or we observe evolutionary branching altogether. These results shed light on how evolution might shape learning mechanisms for social behavior. They suggest that noisy learning does not need to be a by-product of cognitive constraints. Instead, it can serve as a means to gain strategic advantages.

q-bio.PE

Cooperative Dilemmas in Rational Debate

As an epistemic activity, rational debate and discussion requires cooperation, yet involves a tension between collective and individual interests. While all participants benefit from collective outcomes like reaching consensus on true beliefs, individuals face personal costs when changing their minds. This creates an incentive for each debater to let others bear the cognitive burden of exploring alternative perspectives. We present a model to examine the strategic dynamics between debaters motivated by two competing goals: discovering truth and minimizing belief revisions. Our model demonstrates that this tension creates social dilemmas where strategies that are optimal for individuals systematically undermine the collective pursuit of truth. Paradoxically, our analysis reveals that increasing debaters' motivation to seek truth can sometimes produce equilibria with worse outcomes for collective truth discovery. These findings illuminate why rational debate can fail to achieve optimal epistemic outcomes, even when participants genuinely value truth.

cs.SI

Adaptive dynamics of direct reciprocity with N rounds of memory

The theory of direct reciprocity explores how individuals cooperate when they interact repeatedly. In repeated interactions, individuals can condition their behaviour on what happened earlier. One prominent example of a conditional strategy is Tit-for-Tat, which prescribes to cooperate if and only if the co-player did so in the previous round. The evolutionary dynamics among such memory-1 strategies have been explored in quite some detail. However, obtaining analytical results on the dynamics of higher memory strategies becomes increasingly difficult, due to the rapidly growing size of the strategy space. Here, we derive such results for the adaptive dynamics in the donation game. In particular, we prove that for every orbit forward in time, there is an associated orbit backward in time that also solves the differential equation. Moreover, we analyse the dynamics by separating payoffs into a symmetric and an anti-symmetric part and demonstrate some properties of the anti-symmetric part. These results highlight some interesting symmetries that arise when interchanging player one with player two, and cooperation with defection.

math.DS

The co-evolution of direct, indirect and generalized reciprocity

People often engage in costly cooperation, especially in repeated interactions. When deciding whether to cooperate, individuals typically take into account how others have acted in the past. For instance, when one person is deciding whether to cooperate with another, they may consider how they were treated by the other party (direct reciprocity), how the other party treated others (indirect reciprocity), or how they themselves were treated by others in general (generalized reciprocity). Given these different approaches, it is unclear which strategy, or more specifically which mode of reciprocity, individuals will prefer. This study introduces a model where individuals decide how much weight to give each type of information when choosing to cooperate. Through equilibrium analysis, we find that all three modes of reciprocity can be sustained when individuals have sufficiently frequent interactions. However, the existence of such equilibria does not guarantee that individuals will learn to use them. Simulations show that when individuals mainly imitate others, generalized reciprocity often hinders cooperation, leading to defection even under conditions favorable to cooperation. In contrast, when individuals explore new strategies during learning, stable cooperation emerges through direct reciprocity. This study highlights the importance of studying all forms of reciprocity in unison.

q-bio.PE

Indirect reciprocity under opinion synchronization

Indirect reciprocity is a key explanation for the exceptional magnitude of cooperation among humans. This literature suggests that a large proportion of human cooperation is driven by social norms and individuals' incentives to maintain a good reputation. This intuition has been formalized with two types of models. In public assessment models, all community members are assumed to agree on each others' reputations; in private assessment models, people may have disagreements. Both types of models aim to understand the interplay of social norms and cooperation. Yet their results can be vastly different. Public assessment models argue that cooperation can evolve easily, and that the most effective norms tend to be stern. Private assessment models often find cooperation to be unstable, and successful norms show some leniency. Here, we propose a model that can organize these differing results within a single framework. We show that the stability of cooperation depends on a single quantity: the extent to which individual opinions turn out to be correlated. This correlation is determined by a group's norms and the structure of social interactions. In particular, we prove that no cooperative norm is evolutionarily stable when individual opinions are statistically independent. These results have important implications for our understanding of cooperation, conformity, and polarization.

physics.soc-ph

Conditional cooperation with longer memory

Direct reciprocity is a wide-spread mechanism for evolution of cooperation. In repeated interactions, players can condition their behavior on previous outcomes. A well known approach is given by reactive strategies, which respond to the co-player's previous move. Here we extend reactive strategies to longer memories. A reactive-$n$ strategy takes into account the sequence of the last $n$ moves of the co-player. A reactive-$n$ counting strategy records how often the co-player has cooperated during the last $n$ rounds. We derive an algorithm to identify all partner strategies among reactive-$n$ strategies. We give explicit conditions for all partner strategies among reactive-2, reactive-3 strategies, and reactive-$n$ counting strategies. Partner strategies are those that ensure mutual cooperation without exploitation. We perform evolutionary simulations and find that longer memory increases the average cooperation rate for reactive-$n$ strategies but not for reactive counting strategies. Paying attention to the sequence of moves is necessary for reaping the advantages of longer memory.

cs.GT

Evolution of reciprocity with limited payoff memory

Direct reciprocity is a mechanism for the evolution of cooperation in repeated social interactions. According to this literature, individuals naturally learn to adopt conditionally cooperative strategies if they have multiple encounters with their partner. Corresponding models have greatly facilitated our understanding of cooperation, yet they often make strong assumptions on how individuals remember and process payoff information. For example, when strategies are updated through social learning, it is commonly assumed that individuals compare their average payoffs. This would require them to compute (or remember) their payoffs against everyone else in the population. To understand how more realistic constraints influence direct reciprocity, we consider the evolution of conditional behaviors when individuals learn based on more recent experiences. Even in the most extreme case that they only take into account their very last interaction, we find that cooperation can still evolve. However, such individuals adopt less generous strategies, and they tend to cooperate less often than in the classical setup with average payoffs. Interestingly, once individuals remember the payoffs of two or three recent interactions, cooperation rates quickly approach the classical limit. These findings contribute to a literature that explores which kind of cognitive capabilities are required for reciprocal cooperation. While our results suggest that some rudimentary form of payoff memory is necessary, it already suffices to remember a few interactions.

physics.soc-ph

Mutation enhances cooperation in direct reciprocity

Direct reciprocity is a powerful mechanism for evolution of cooperation based on repeated interactions between the same individuals. But high levels of cooperation evolve only if the benefit-to-cost ratio exceeds a certain threshold that depends on memory length. For the best-explored case of one-round memory, that threshold is two. Here we report that intermediate mutation rates lead to high levels of cooperation, even if the benefit-to-cost ratio is only marginally above one, and even if individuals only use a minimum of past information. This surprising observation is caused by two effects. First, mutation generates diversity which undermines the evolutionary stability of defectors. Second, mutation leads to diverse communities of cooperators that are more resilient than homogeneous ones. This finding is relevant because many real world opportunities for cooperation have small benefit-to-cost ratios, which are between one and two, and we describe how direct reciprocity can attain cooperation in such settings. Our result can be interpreted as showing that diversity, rather than uniformity, promotes evolution of cooperation.

q-bio.PE

Indirect reciprocity with stochastic rules

Cooperation is a crucial aspect of social life, yet understanding the nature of cooperation and how it can be promoted is an ongoing challenge. One mechanism for cooperation is indirect reciprocity. According to this mechanism, individuals cooperate to maintain a good reputation. This idea is embodied in a set of social norms called the ``leading eight''. When all information is publicly available, these norms have two major properties. Populations that employ these norms are fully cooperative, and they are stable against invasion by alternative norms. In this paper, we extend the framework of the leading eight in two directions. First, we allow social norms to be stochastic. Such norms allow individuals to evaluate others with certain probabilities. Second, we consider norms in which also the reputations of passive recipients can be updated. Using this framework, we characterize all evolutionarily stable norms that lead to full cooperation in the public information regime. When only the donor's reputation is updated, and all updates are deterministic, we recover the conventional model. In that case, we find two classes of stable norms: the leading eight and the `secondary sixteen'. Stochasticity can further help to stabilize cooperation when the benefit of cooperation is comparably small. Moreover, updating the recipients' reputations can help populations to recover more quickly from errors. Overall, our study highlights a remarkable trade-off between the evolutionary stability of a norm and its robustness with respect to errors. Norms that correct errors quickly require higher benefits of cooperation to be stable.

q-bio.PE

Evolutionary instability of selfish learning in repeated games

Across many domains of interaction, both natural and artificial, individuals use past experience to shape future behaviors. The results of such learning processes depend on what individuals wish to maximize. A natural objective is one's own success. However, when two such "selfish" learners interact with each other, the outcome can be detrimental to both, especially when there are conflicts of interest. Here, we explore how a learner can align incentives with a selfish opponent. Moreover, we consider the dynamics that arise when learning rules themselves are subject to evolutionary pressure. By combining extensive simulations and analytical techniques, we demonstrate that selfish learning is unstable in most classical two-player repeated games. If evolution operates on the level of long-run payoffs, selection instead favors learning rules that incorporate social (other-regarding) preferences. To further corroborate these results, we analyze data from a repeated prisoner's dilemma experiment. We find that selfish learning is insufficient to explain human behavior when there is a trade-off between payoff maximization and fairness.

q-bio.PE

Evolution of direct reciprocity in group-structured populations

People tend to have their social interactions with members of their own community. Such group-structured interactions can have a profound impact on the behaviors that evolve. Group structure affects the way people cooperate, and how they reciprocate each other's cooperative actions. Past work has shown that population structure and reciprocity can both promote the evolution of cooperation. Yet the impact of these mechanisms has been typically studied in isolation. In this work, we study how the two mechanisms interact. Using a game-theoretic model, we explore how people engage in reciprocal cooperation in group-structured populations, compared to well-mixed populations of equal size. To derive analytical results, we focus on two scenarios. In the first scenario, we assume a complete separation of time scales. Mutations are rare compared to between-group comparisons, which themselves are rare compared to within-group comparisons. In the second scenario, there is a partial separation of time scales, where mutations and between-group comparisons occur at a comparable rate. In both scenarios, we find that the effect of population structure depends on the benefit of cooperation. When this benefit is small, group-structured populations are more cooperative. But when the benefit is large, well-mixed populations result in more cooperation. Overall, our results reveal how group structure can sometimes enhance and sometimes suppress the evolution of cooperation.

q-bio.PE

Comparing reactive and memory-one strategies of direct reciprocity

Direct reciprocity is a mechanism for the evolution of cooperation based on repeated interactions. When individuals meet repeatedly, they can use conditional strategies to enforce cooperative outcomes that would not be feasible in one-shot social dilemmas. Direct reciprocity requires that individuals keep track of their past interactions and find the right response. However, there are natural bounds on strategic complexity: Humans find it difficult to remember past interactions accurately, especially over long timespans. Given these limitations, it is natural to ask how complex strategies need to be for cooperation to evolve. Here, we study stochastic evolutionary game dynamics in finite populations to systematically compare the evolutionary performance of reactive strategies, which only respond to the co-player's previous move, and memory-one strategies, which take into account the own and the co-player's previous move. In both cases, we compare deterministic strategy and stochastic strategy spaces. For reactive strategies and small costs, we find that stochasticity benefits cooperation, because it allows for generous-tit-for-tat. For memory one strategies and small costs, we find that stochasticity does not increase the propensity for cooperation, because the deterministic rule of win-stay, lose-shift works best. For memory one strategies and large costs, however, stochasticity can augment cooperation.

q-bio.PE

Zero-determinant alliances in multiplayer social dilemmas

Direct reciprocity and conditional cooperation are important mechanisms to prevent free riding in social dilemmas. But in large groups these mechanisms may become ineffective, because they require single individuals to have a substantial influence on their peers. However, the recent discovery of the powerful class of zero-determinant strategies in the iterated prisoner's dilemma suggests that we may have underestimated the degree of control that a single player can exert. Here, we develop a theory for zero-determinant strategies for multiplayer social dilemmas, with any number of involved players. We distinguish several particularly interesting subclasses of strategies: fair strategies ensure that the own payoff matches the average payoff of the group; extortionate strategies allow a player to perform above average; and generous strategies let a player perform below average. We use this theory to explore how individuals can enhance their strategic options by forming alliances. The effects of an alliance depend on the size of the alliance, the type of the social dilemma, and on the strategy of the allies: fair alliances reduce the inequality within their group; extortionate alliances outperform the remaining group members; but generous alliances increase welfare. Our results highlight the critical interplay of individual control and alliance formation to succeed in large groups.

q-bio.PE