SearcharxivSearch

arXiv subjects

Yian Yin

Publications and source records attributed to Yian Yin.

At least 19 recordsLinked to original sources

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.

cs.DL

Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment

Collaboration is the defining mode of modern science, yet its core mechanism -- feedback -- remains hard to observe, difficult to scale, and unequally distributed. Here we test whether large language models (LLMs) can contribute to this hidden but vital practice and reallocate scientific feedback, an essential yet scarce resource for knowledge production. In a global large-scale randomized field experiment, we delivered customized LLM-generated feedback for over 31,000 arXiv preprints across 150 fields and more than 45,000 researchers from 133 geographic regions. Relative to controls, authors who received feedback had a significantly higher likelihood of revising their manuscripts, corresponding to a 12.55% relative increase over the baseline revision rate. Exposure to AI feedback also increased authors' subsequent use of LLM tools in their future papers, suggesting longer-run shifts in scientific practice. These effects were strongest among authors from non-English-dominant research regions, manuscripts less embedded in the scholarly literature, and teams with lower h-indexes and earlier career stages, consistent with the idea that AI feedback may provide the greatest benefit where access to timely critique is otherwise limited. Together, these findings provide causal evidence that structured AI-based interventions can transform access to scientific feedback from a largely private advantage into a more widely distributed resource, with broader implications for productivity, equity, and capacity across the global research system.

physics.soc-ph

Universal Dynamics of Punctuated Progress

Scientific and technological frontiers advance through punctuated dynamics, yet the principles governing these dynamics remain poorly understood. Here we collect and analyze datasets tracking the evolution of frontiers across 9 different domains, spanning materials discovery, structural biology, AI, computational biomedicine, data science, theoretical computer science, Formula-1 racing, and physical wheel building. Analyzing 6.8M solutions to 6.7K tasks, we uncover three universal patterns: (1) waiting times between new frontiers are heavy-tailed, with most attempts concentrated in long stasis; (2) frontier records accumulate at a sublinear rate, faster than logarithmic yet slower than linear growth; (3) record-breaking events are temporally correlated, generating short-term predictability yet long-term unpredictability. Despite the differences in the scale, scope, and definition of the settings, these patterns are remarkably consistent across all domains we study, and are not captured by models from complex systems, record statistics, economics of innovation, and cultural evolution. We trace the missing ingredient to the distinction between radical and incremental innovation, and develop a minimal, analytically solvable model incorporating both radical resets that restructure what is achievable and incremental refinements that exploit the current frontier. The simple model reproduces all three empirical regularities. Remarkably, the leading-order predictions are parameter-independent, identifying a new universality class governing punctuated progress and yielding testable predictions about how openness and access to frontier solutions shape the pace of advance. Overall, these results reveal universal dynamics governing punctuated progress and identify the interplay between radical resets and incremental refinements as the key driver of how scientific and technological frontiers advance.

physics.soc-ph

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a uniquely verifiable object - scientific citations - to audit 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central. We find a sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. These errors are diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams. At the same time, hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition. Preprint moderation and journal publication processes capture only a fraction of these errors, suggesting that the spread of hallucinated content has outpaced existing safeguards. Together, these findings demonstrate that LLM hallucinations are infiltrating knowledge production at scale, threatening both the reliability and equity of future scientific discovery as human and AI systems draw on the existing literature.

cs.DL

Scientific production in the era of Large Language Models

Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.

cs.DL

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years later, by higher impact like citations, sales, or reviews. We formalize this setting as Lead-Lag Forecasting (LLF): given an early usage channel (the lead), predict a correlated but temporally shifted outcome channel (the lag). Despite the ubiquity of such patterns, LLF has not been treated as a unified forecasting problem within the time-series community, largely due to the absence of standardised datasets. To anchor research in LLF, here we present two high-volume benchmark datasets: arXiv (accesses -> citations of 2.3M papers) and GitHub (pushes/stars -> forks of 3M repositories). Our datasets provide ideal testbeds for lead-lag forecasting, by capturing long-horizon dynamics across years, spanning the full spectrum of outcomes, and avoiding survivorship bias in sampling. We documented all technical details of data curation and cleaning, verified the presence of lead-lag dynamics through statistical and classification tests, and benchmarked parametric and non-parametric baselines for regression. Our study establishes LLF as a novel forecasting paradigm and lays an empirical foundation for its systematic exploration in social and usage data.

cs.LG

Survivors, Complainers, and Borderliners: Upward Bias in Online Discussions of Academic Conference Reviews

Online discussion platforms, such as community Q&A sites and forums, have become important hubs where academic conference authors share and seek information about the peer review process and outcomes. However, these discussions involve only a subset of all submissions, raising concerns about the representativeness of the self-reported review scores. In this paper, we conduct a systematic study comparing the review score distributions of self-reported submissions in online discussions (based on data collected from Zhihu and Reddit) with those of all submissions. We reveal a consistent upward bias: the score distribution of self-reported samples is shifted upward relative to the population score distribution, with this difference statistically significant in most cases. Our analysis identifies three distinct contributors to this bias: (1) survivors, authors of accepted papers who are more likely to share good results than those of rejected papers who tend to conceal bad ones; (2) complainers, authors of high-scoring rejected papers who are more likely to voice complaints about the peer review process or outcomes than those of low scores; and (3) borderliners, authors with borderline scores who face greater uncertainty prior to decision announcements and are more likely to seek advice during the rebuttal period. These findings have important implications for how information seekers should interpret online discussions of academic conference reviews.

cs.SI

Funding the Frontier: Visualizing the Broad Impact of Science and Science Funding

Understanding the broad impact of science and science funding is critical to ensuring that science investments and policies align with societal needs. Existing research links science funding to the output of scientific publications but largely leaves out the downstream uses of science and the myriad ways in which investing in science may impact human society. As funders seek to allocate scarce funding resources across a complex research landscape, there is an urgent need for informative and transparent tools that allow for comprehensive assessments and visualization of the impact of funding. Here we present Funding the Frontier (FtF), a visual analysis system for researchers, funders, policymakers, university leaders, and the broad public to analyze multidimensional impacts of funding and make informed decisions regarding research investments and opportunities. The system is built on a massive data collection that connects 7M research grants to 140M scientific publications, 160M patents, 10.9M policy documents, 800K clinical trials, and 5.8M newsfeeds, with 1.8B citation linkages among these entities, systematically linking science funding to its downstream impacts. As such, Funding the Frontier is distinguished by its multifaceted impact analysis framework. The system incorporates diverse impact metrics and predictive models that forecast future investment opportunities into an array of coordinated views, allowing for easy exploration of funding and its outcomes. We evaluate the effectiveness and usability of the system using case studies and expert interviews. Feedback suggests that our system not only fulfills the primary analysis needs of its target users, but the rich datasets of the complex science ecosystem and the proposed analysis framework also open new avenues for both visualization and the science of science research.

cs.HC

Synthesis of innovation and obsolescence

Innovation and obsolescence describe the dynamics of ever-churning social and biological systems, from the development of economic markets to scientific and technological progress to biological evolution. They have been widely discussed, but in isolation, leading to fragmented modeling of their dynamics. This poses a problem for connecting and building on what we know about their shared mechanisms. Here we collectively propose a conceptual and mathematical framework to transcend field boundaries and to explore unifying theoretical frameworks and open challenges. We ring an optimistic note for weaving together disparate threads with key ideas from the wide and largely disconnected literature by focusing on the duality of innovation and obsolescence and by proposing a mathematical framework to unify the metaphors between constitutive elements.

physics.soc-ph

Adaptability and the Pivot Penalty in Science and Technology

Scientists and inventors set the direction of their work amidst an evolving landscape of questions, opportunities, and challenges. This paper introduces a measurement framework to quantify how far researchers move from their existing research when producing new works. We apply this framework to millions of scientific publications and patents and uncover a pervasive "pivot penalty", where the impact of new research steeply declines the further a researcher moves from their prior work. The pivot penalty applies nearly universally across scientific publishing and patenting and has been growing in magnitude over the past five decades. While creativity frameworks suggest a benefit to exploratory search by researchers and often emphasize outsider advantages in driving breakthroughs, we find little evidence for such an advantage. The pivot penalty is consistent with increasingly narrow specializations of researchers, and when researchers undertake large pivots, a signature of their work is weak engagement with established mixtures of prior knowledge. Unexpected shocks to the research landscape, which may push researchers away from existing areas or pull them into new ones, further demonstrate substantial pivot penalties. COVID-19 provides a high-scale case study, where many researchers engaged the pandemic, yet the pivot penalty remains severe. The pivot penalty generalizes across fields, career stage, productivity, collaboration, and funding contexts, highlighting both the breadth and depth of the adaptive challenge. Overall, the findings point to large and increasing challenges in adapting to new opportunities and threats. The results have implications for individual researchers, research organizations, science policy, and the capacity of science and society as a whole to confront emergent demands.

cs.DL

Language Agents for Detecting Implicit Stereotypes in Text-to-image Models at Scale

The recent surge in the research of diffusion models has accelerated the adoption of text-to-image models in various Artificial Intelligence Generated Content (AIGC) commercial products. While these exceptional AIGC products are gaining increasing recognition and sparking enthusiasm among consumers, the questions regarding whether, when, and how these models might unintentionally reinforce existing societal stereotypes remain largely unaddressed. Motivated by recent advancements in language agents, here we introduce a novel agent architecture tailored for stereotype detection in text-to-image models. This versatile agent architecture is capable of accommodating free-form detection tasks and can autonomously invoke various tools to facilitate the entire process, from generating corresponding instructions and images, to detecting stereotypes. We build the stereotype-relevant benchmark based on multiple open-text datasets, and apply this architecture to commercial products and popular open source text-to-image models. We find that these models often display serious stereotypes when it comes to certain prompts about personal characteristics, social cultural context and crime-related aspects. In summary, these empirical findings underscore the pervasive existence of stereotypes across social dimensions, including gender, race, and religion, which not only validate the effectiveness of our proposed approach, but also emphasize the critical necessity of addressing potential ethical risks in the burgeoning realm of AIGC. As AIGC continues its rapid expansion trajectory, with new models and plugins emerging daily in staggering numbers, the challenge lies in the timely detection and mitigation of potential biases within these models.

cs.CY

Can large language models provide useful feedback on research papers? A large-scale empirical analysis

Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4's feedback through two large-scale studies. We first quantitatively compared GPT-4's generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers. We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations.

cs.LG

Loss of New Ideas: Potentially Long-lasting Effects of the Pandemic on Scientists

Extensive research has documented the immediate impacts of the COVID-19 pandemic on scientists, yet it remains unclear if and how such impacts have shifted over time. Here we compare results from two surveys of principal investigators, conducted between April 2020 and January 2021, along with analyses of large-scale publication data. We find that there has been a clear sign of recovery in some regards, as scientists' time spent on their work has almost returned to pre-pandemic levels. However, the latest data also reveals a new dimension in which the pandemic is affecting the scientific workforce: the rate of initiating new research projects. Except for the small fraction of scientists who directly engaged in COVID-related research, most scientists started significantly fewer new research projects in 2020. This decline is most pronounced amongst the same demographic groups of scientists who reported the largest initial disruptions: female scientists and those with young children. Yet in sharp contrast to the earlier phase of the pandemic, when there were large disparities across scientific fields, this loss of new projects appears remarkably homogeneous across fields. Analyses of large-scale publication data reveal a global decline in the rate of new collaborations, especially in non-COVID-related preprints, which is consistent with the reported decline in new projects. Overall, these findings highlight that, while the end of the pandemic may appear in sight in some countries, its large and unequal impact on the scientific workforce may be enduring, which may have broad implications for inequality and the long-term vitality of science.

cs.DL

Science as a Public Good: Public Use and Funding of Science

Knowledge of how science is consumed in public domains is essential for a deeper understanding of the role of science in human society. While science is heavily supported by public funding, common depictions suggest that scientific research remains an isolated or 'ivory tower' activity, with weak connectivity to public use, little relationship between the quality of research and its public use, and little correspondence between the funding of science and its public use. This paper introduces a measurement framework to examine public good features of science, allowing us to study public uses of science, the public funding of science, and how use and funding relate. Specifically, we integrate five large-scale datasets that link scientific publications from all scientific fields to their upstream funding support and downstream public uses across three public domains - government documents, the news media, and marketplace invention. We find that the public uses of science are extremely diverse, with different public domains drawing distinctively across scientific fields. Yet amidst these differences, we find key forms of alignment in the interface between science and society. First, despite concerns that the public does not engage high-quality science, we find universal alignment, in each scientific field and public domain, between what the public consumes and what is highly impactful within science. Second, despite myriad factors underpinning the public funding of science, the resulting allocation across fields presents a striking alignment with the field's collective public use. Overall, public uses of science present a rich landscape of specialized consumption, yet collectively science and society interface with remarkable, quantifiable alignment between scientific use, public use, and funding.

cs.DL

Quantifying Policy Responses to a Global Emergency: Insights from the COVID-19 Pandemic

Public policy must confront emergencies that evolve in real time and in uncertain directions, yet little is known about the nature of policy response. Here we take the coronavirus pandemic as a global and extraordinarily consequential case, and study the global policy response by analyzing a novel dataset recording policy documents published by government agencies, think tanks, and intergovernmental organizations (IGOs) across 114 countries (37,725 policy documents from Jan 2nd through May 26th 2020). Our analyses reveal four primary findings. (1) Global policy attention to COVID-19 follows a remarkably similar trajectory as the total confirmed cases of COVID-19, yet with evolving policy focus from public health to broader social issues. (2) The COVID-19 policy frontier disproportionately draws on the latest, peer-reviewed, and high-impact scientific insights. Moreover, policy documents that cite science appear especially impactful within the policy domain. (3) The global policy frontier is primarily interconnected through IGOs, such as the WHO, which produce policy documents that are central to the COVID19 policy network and draw especially strongly on scientific literature. Removing IGOs' contributions fundamentally alters the global policy landscape, with the policy citation network among government agencies increasingly fragmented into many isolated clusters. (4) Countries exhibit highly heterogeneous policy attention to COVID-19. Most strikingly, a country's early policy attention to COVID-19 shows a surprising degree of predictability for the country's subsequent deaths. Overall, these results uncover fundamental patterns of policy interactions and, given the consequential nature of emergent threats and the paucity of quantitative approaches to understand them, open up novel dimensions for assessing and effectively coordinating global and local responses to COVID-19 and beyond.

physics.soc-ph

Quantifying the Immediate Effects of the COVID-19 Pandemic on Scientists

The COVID-19 pandemic has undoubtedly disrupted the scientific enterprise, but we lack empirical evidence on the nature and magnitude of these disruptions. Here we report the results of a survey of approximately 4,500 Principal Investigators (PIs) at U.S.- and Europe-based research institutions. Distributed in mid-April 2020, the survey solicited information about how scientists' work changed from the onset of the pandemic, how their research output might be affected in the near future, and a wide range of individuals' characteristics. Scientists report a sharp decline in time spent on research on average, but there is substantial heterogeneity with a significant share reporting no change or even increases. Some of this heterogeneity is due to field-specific differences, with laboratory-based fields being the most negatively affected, and some is due to gender, with female scientists reporting larger declines. However, among the individuals' characteristics examined, the largest disruptions are connected to a usually unobserved dimension: childcare. Reporting a young dependent is associated with declines similar in magnitude to those reported by the laboratory-based fields and can account for a significant fraction of gender differences. Amidst scarce evidence about the role of parenting in scientists' work, these results highlight the fundamental and heterogeneous ways this pandemic is affecting the scientific workforce, and may have broad relevance for shaping responses to the pandemic's effect on science and beyond.

physics.soc-ph

Scientific elite revisited: Patterns of productivity, collaboration, authorship and impact

Throughout history, a relatively small number of individuals have made a profound and lasting impact on science and society. Despite long-standing, multi-disciplinary interests in understanding careers of elite scientists, there have been limited attempts for a quantitative, career-level analysis. Here, we leverage a comprehensive dataset we assembled, allowing us to trace the entire career histories of nearly all Nobel laureates in physics, chemistry, and physiology or medicine over the past century. We find that, although Nobel laureates were energetic producers from the outset, producing works that garner unusually high impact, their careers before winning the prize follow relatively similar patterns as ordinary scientists, being characterized by hot streaks and increasing reliance on collaborations. We also uncovered notable variations along their careers, often associated with the Nobel prize, including shifting coauthorship structure in the prize-winning work, and a significant but temporary dip in the impact of work they produce after winning the Nobel. Together, these results document quantitative patterns governing the careers of scientific elites, offering an empirical basis for a deeper understanding of the hallmarks of exceptional careers in science.

cs.DL

Quantifying dynamics of failure across science, startups, and security

Human achievements are often preceded by repeated attempts that initially fail, yet little is known about the mechanisms governing the dynamics of failure. Here, building on the rich literature on innovation, human dynamics and learning, we develop a simple one-parameter model that mimics how successful future attempts build on those past. Analytically solving this model reveals a phase transition that separates dynamics of failure into regions of stagnation or progression, predicting that near the critical threshold, agents who share similar characteristics and learning strategies may experience fundamentally different outcomes following failures. Below the critical point, we see those who explore disjoint opportunities without a pattern of improvement, and above it, those who exploit incremental refinements to systematically advance toward success. The model makes several empirically testable predictions, demonstrating that those who eventually succeed and those who do not may be initially similar, yet are characterized by fundamentally distinct failure dynamics in terms of the efficiency and quality of each subsequent attempt. We collected large-scale data from three disparate domains, tracing repeated attempts by (i) NIH investigators to fund their research, (ii) innovators to successfully exit their startup ventures, and (iii) terrorist organizations to post casualties in violent attacks, finding broadly consistent empirical support across all three domains. Together, our findings unveil identifiable yet previously unknown early signals that allow us to identify failure dynamics that will lead to ultimate victory or defeat. Given the ubiquitous nature of failures and the paucity of quantitative approaches to understand them, these results represent a crucial step toward deeper understanding of the complex dynamics beneath failures, the essential prerequisites for success.

physics.soc-ph