SearcharxivSearch

arXiv subjects

Linzhuo Li

Publications and source records attributed to Linzhuo Li.

12 recordsLinked to original sources

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities

As LLMs become increasingly integrated into human society, evaluating their orientations on human values from social science has drawn growing attention. Nevertheless, it is still unclear why human values matter for LLMs, especially in LLM-based multi-agent systems, where group-level failures may accumulate from individually misaligned actions. We ask whether misalignment with human values alters the collective behavior of LLM agents and what changes it induces? In this work, we introduce CIVA, a controlled multi-agent environment grounded in social science theories, where LLM agents form a community and autonomously communicate, explore, and compete for resources, enabling systematic manipulation of value prevalence and behavioral analysis. Through comprehensive simulation experiments, we reveal three key findings. (1) We identify several structurally critical values that substantially shape the community's collective dynamics, including those diverging from LLMs' original orientations. Triggered by the misspecification of these values, we (2) detect system failure modes, e.g., catastrophic collapse, at the macro level, and (3) observe emergent behaviors like deception and power-seeking at the micro level. These results offer quantitative evidence that human values are essential for collective outcomes in LLMs and motivate future multi-agent value alignment.

cs.CL

Knowing Your Uncertainty -- On the application of LLM in social sciences

Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique challenges for scientific inquiry. This article argues that applying LLMs to social scientific tasks requires explicit assessment of uncertainty -- an expectation long established in both quantitative methodology in the social sciences and machine learning. We introduce a unified framework for evaluating LLM uncertainty based on Hill numbers, a family of diversity measures. By transforming existing uncertainty quantification (UQ) metrics into Hill numbers, the framework provides a common and intuitive scale for interpreting variation in LLM outputs while accommodating different notions of semantic similarity and different sensitivities to output distributions. We show how it might help the application of LLMs in social sciences through four empirical applications.

cs.CY

Innovation by Displacement

New ideas are often thought to arise from recombining existing knowledge. Yet despite rapid publication growth - and expanding opportunities for recombination - scientific breakthroughs remain rare. This gap between productivity and progress challenges recombinant growth theory as the prevailing account of innovation. We argue that the limitation of this theory lies in treating ideas solely as complements, overlooking that breakthroughs often arise when ideas act as substitutes. To test this, we integrate scientist interviews, bibliometric validation, and machine learning analysis of 41 million papers (1965 - 2024). Interviews reveal that breakthroughs are marked not by novelty (Atypicality) alone but by their ability to displace dominant ideas (Disruption). Large-scale analysis confirms that novelty and disruption represent distinct innovation mechanisms: they are negatively correlated across domains, periods, team sizes, and paper versions. Novel papers extend dominant ideas across topics and attract immediate attention; disruptive papers displace them within the same topic and generate lasting influence. Hence, progress slows not from lack of effort but because most research extends rather than overturns ideas. Applying this perspective reveals distinct roles of theories and methods in scientific change: methods more often drive breakthroughs, whereas theories tend to be novel but rarely disruptive, reinforcing the dominance of established ideas.

cs.DL

ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models

Large language models (LLMs) in the form of chatbots like ChatGPT and Llama are increasingly proposed as "silicon samples" for simulating human opinions. This study examines this notion, arguing that LLMs may misrepresent population-level opinions. We identify two fundamental challenges: a failure in structural consistency, where response accuracy doesn't hold across demographic aggregation levels, and homogenization, an underrepresentation of minority opinions. To investigate these, we prompted ChatGPT (GPT-4) and Meta's Llama 3.1 series (8B, 70B, 405B) with questions on abortion and unauthorized immigration from the American National Election Studies (ANES) 2020. Our findings reveal significant structural inconsistencies and severe homogenization in LLM responses compared to human data. We propose an "accuracy-optimization hypothesis," suggesting homogenization stems from prioritizing modal responses. These issues challenge the validity of using LLMs, especially chatbots AI, as direct substitutes for human survey data, potentially reinforcing stereotypes and misinforming policy.

cs.CL

Can Recombination Displace Dominant Scientific Ideas

The combination of diverse, pre-existing knowledge is a common explanation for scientific breakthroughs. However, a paradox exists: while scientific output and the potential for such recombination have grown exponentially, the rate of breakthrough discoveries has not. To explore this paradox, our study examines 41 million scientific articles from 1965 to 2024. We measure two key properties for each paper: atypicality, which quantifies the combination of knowledge from conceptually distant areas, and disruption. We demonstrate that these metrics capture distinct processes. Atypicality is characteristic of work that extends established concepts into new topical areas (a form of cross-topic recombination). Disruption, in contrast, signifies the replacement of a dominant idea within its own topic.

cs.DL

Tinkering Against Scaling

The ascent of scaling in artificial intelligence research has revolutionized the field over the past decade, yet it presents significant challenges for academic researchers, particularly in computational social science and critical algorithm studies. The dominance of large language models, characterized by their extensive parameters and costly training processes, creates a disparity where only industry-affiliated researchers can access these resources. This imbalance restricts academic researchers from fully understanding their tools, leading to issues like reproducibility in computational social science and a reliance on black-box metaphors in critical studies. To address these challenges, we propose a "tinkering" approach that is inspired by existing works. This method involves engaging with smaller models or components that are manageable for ordinary researchers, fostering hands-on interaction with algorithms. We argue that tinkering is both a way of making and knowing for computational social science and a way of knowing for critical studies, and fundamentally, it is a way of caring that has broader implications for both fields.

cs.CY

The Disruption Index Measures Displacement Between a Paper and Its Most Cited Reference

Initially developed to capture technical innovation and later adapted to identify scientific breakthroughs, the Disruption Index (D-index) offers the first quantitative framework for analyzing transformative research. Despite its promise, prior studies have struggled to clarify its theoretical foundations, raising concerns about potential bias. Here, we show that-contrary to the common belief that the D-index measures absolute innovation-it captures relative innovation: a paper's ability to displace its most-cited reference. In this way, the D-index reflects scientific progress as the replacement of older answers with newer ones to the same fundamental question-much like light bulbs replacing candles. We support this insight through mathematical analysis, expert surveys, and large-scale bibliometric evidence. To facilitate replication, validation, and broader use, we release a dataset of D-index values for 49 million journal articles (1800-2024) based on OpenAlex.

cs.DL

Is Science Inevitable?

Using large-scale citation data and a breakthrough metric, the study systematically evaluates the inevitability of scientific breakthroughs. We find that scientific breakthroughs emerge as multiple discoveries rather than singular events. Through analysis of over 40 million journal articles, we identify multiple discoveries as papers that independently displace the same reference using the Disruption Index (D-index), suggesting functional equivalence. Our findings support Merton's core argument that scientific discoveries arise from historical context rather than individual genius. The results reveal a long-tail distribution pattern of multiple discoveries across various datasets, challenging Merton's Poisson model while reinforcing the structural inevitability of scientific progress.

cs.DL

Team Size and Its Negative Impact on the Disruption Index

As science transitions from the age of lone geniuses to an era of collaborative teams, the question of whether large teams can sustain the creativity of individuals and continue driving innovation has become increasingly important. Our previous research first revealed a negative relationship between team size and the Disruption Index-a network-based metric of innovation-by analyzing 65 million projects across papers, patents, and software over half a century. This work has sparked lively debates within the scientific community about the robustness of the Disruption Index in capturing the impact of team size on innovation. Here, we present additional evidence that the negative link between team size and disruption holds, even when accounting for factors such as reference length, citation impact, and historical time. We further show how a narrow 5-year window for measuring disruption can misrepresent this relationship as positive, underestimating the long-term disruptive potential of small teams. Like "sleeping beauties," small teams need a decade or more to see their transformative contributions to science.

cs.SI

Displacing Science

Recent research on the decline in the paper disruption index (D-index) has sparked heated debates among scholars and garnered significant attention from policymakers and research institution leaders globally. To bridge the gap between policymakers' interest and scholars' skepticism about the D-index, we present this article summarizing key insights from our eight-year investigation, including interviews with scientists across nine disciplines and an analysis of 41 million papers over six decades. Our work confirms the decline in disruptive papers, addresses relevant technical concerns, and makes several original contributions: we clarify that the D-index measures how new ideas render old ones obsolete, suggesting 'Displacing' as an alternative interpretation for 'D'; we show that federal funding agencies like the NIH and NSF are less likely to support disruptive research; and we introduce the 'principle of functional equivalence' to explain the origins of recombining and displacing mechanisms in science, stressing that not all innovation problems are combinatorial, and challenging the belief that AI is the mature solution to scientific innovation. This article aims to promote broader and more accurate use of the D-index in research evaluation and to inspire new funding mechanisms for scientific breakthroughs.

cs.DL

Social Centralization and Semantic Collapse: Hyperbolic Embeddings of Networks and Text

Modern advances in transportation and communication technology from airplanes to the internet alongside global expansions of media, migration, and trade have made the modern world more connected than ever before. But what does this bode for the convergence of global culture? Here we explore the relationship between centralization in social networks and contraction or collapse in the diversity of semantic expressions such as ideas, opinions, and tastes. We advance formal examination of this relationship by introducing new methods of manifold learning that allow us to map social networks and semantic combinations into comparable hyperbolic spaces. Hyperbolic representations natively represent both hierarchy and diversity within a system. We illustrate this method by examining the relationship between social centralization and semantic diversity within 21st Century physics, empirically demonstrating how dense, centralized collaboration is associated with a reduction in the space of ideas and how these patterns generalize to all modern scholarship and science. We discuss the complex of causes underlying this association, and theorize the dynamic interplay between structural centralization and semantic contraction, arguing that it introduces an essential tension between the supply and demand of difference.

cs.SI

Social Connection Induces Cultural Contraction: Evidence from Hyperbolic Embeddings of Social and Semantic Networks

Research has repeatedly demonstrated the influence of social connection and communication on convergence in cultural tastes, opinions and ideas. Here we review recent studies and consider the implications of social connection on cultural, epistemological and ideological contraction, then formalize these intuitions within the language of information theory. To systematically examine connectivity and cultural diversity, we introduce new methods of manifold learning to map both social networks and topic combinations into comparable, two-dimensional hyperbolic spaces or Poincaré disks, which represent both hierarchy and diversity within a system. On a Poincaré disk, radius from center traces the position of an actor in a social hierarchy or an idea in a cultural hierarchy. The angle of the disk required to inscribe connected actors or ideas captures their diversity. Using this method in the epistemic culture of 21st Century physics, we empirically demonstrate that denser university collaborations systematically contract the space of topics discussed and ideas investigated more than shared topics drive collaboration, despite the extreme commitments academic physicists make to research programs over the course of their careers. Dense connections unleash flows of communication that contract otherwise fragmented semantic spaces into convergent hubs or polarized clusters. We theorize the dynamic interplay between structural expansion and cultural contraction and explore how this introduces an essential tension between the enjoyment and protection of difference.

physics.soc-ph