arXiv · 2602.16662
Evaluating Collective Behaviour of Hundreds of LLM Agents
Abstract
LLM-powered AI assistants acting on behalf of users can produce poor collective outcomes at scale. We introduce a framework for evaluating their emergent behaviour in social dilemmas, applied to three iterated games (Public Goods, Collective Risk, Common Pool Resource). We prompt each model to produce a natural-language strategy, then have the same model translate it into code. This aims to isolate strategic reasoning from input-parsing, enables pre-deployment inspection, and scales to populations of hundreds of agents. We propose three analyses: behavioural fingerprinting via exhaustive evaluation over opponent histories; self-play robustness across mixtures of a model's strategies with either a Selfish or Collective disposition; and cultural evolution under payoff-biased imitation. Applied to three state-of-the-art LLMs, we find substantial cross-model differences in self-play welfare, and that cultural evolution converges to low-welfare, Selfish-dominant equilibria in larger groups.
Explore related subjects
Keep this discovery
Richard Willis, Jianing Zhao, Yali Du, Joel Z. Leibo. 2026-02-18. Evaluating Collective Behaviour of Hundreds of LLM Agents. https://arxiv.org/abs/2602.16662
Cite the original work for its findings. Save a collection to share your selection of sources.