Searcharxiv⌕ Search

arXiv · 2609.28504

AI in Science: Early Insights

Abstract

Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements-- LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri, Arthur Turrell, Julian Jacobs, Atoosa Kasirzadeh, Ana Trisovic, Yiyuan Chen, Tanya Rodchenko, Catherine Pollard, Scott Strand, Daniel Rock, Zanna Iscenko, Fabien Curto Millet, Neil Thompson, James Manyika. 2026-09-25. AI in Science: Early Insights. https://arxiv.org/abs/2609.28504

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Delphos: A reinforcement learning framework for assisting discrete choice model specification

We introduce Delphos, a deep reinforcement learning framework for assisting discrete choice model specification process. Delphos aims to support the modeller by providing automated, data-driven suggestions for model specifications, thereby reducing the effort required to develop and refine utility functions. Delphos conceptualises model specification as a sequential decision-making problem, inspired by the way human choice modellers iteratively construct models through a series of reasoned specification decisions. In this setting, an agent learns to specify candidate model specifications by choosing a sequence of modelling actions, such as adding alternative specific constants, accommodating both generic and alternative-specific taste parameters, applying non-linear transformations to attributes, and including interactions with covariates. Each resulting candidate model is estimated and evaluated using a reward function defined by the modeller, which can reflect statistical model fit as well as behavioural expectations. Specifically, Delphos uses a Deep Q-Network to learn how individual specification decisions contribute to the eventual quality of the resulting model and, in turn, which sequences of modelling decisions tend to produce well-performing candidates. We evaluate Delphos on both simulated and empirical datasets using alternative reward functions. In simulated cases, learning curves, Q-value patterns, and performance metrics show that Delphos learns effective specification strategies while exploring only a small fraction of the feasible modelling space. We further apply the framework to two empirical datasets to benchmark and demonstrate its practical use. These experiments illustrate the ability of Delphos to generate competitive, behaviourally plausible models and highlight the potential of this adaptive, learning-based framework to assist the model specification process.

econ.GN↗

Risk Ceilings and Development Deadlines: Pacing AI under Uncertain Safety Productivity

Can a regulator promise both a risk ceiling and a development deadline when safety productivity is unknown? A ceiling below the final model's unprotected hazard requires a minimum stock of safety knowledge, so both promises hold only if weak research can be ruled out. Learning first and then replaying development, with spare compute in safety, comes close to that minimum. In a calibration with only state risk, constant safety yield, and full knowledge transfer, it needs only 3.3 percent more productivity than the necessary bound. Customer services pay for the guarantee. Rules that fix compute allocation fix dates, not risk.

econ.GN↗

Who Leads and Who Collects:Algorithmic Collusion in Markets of Heterogeneous Language Models

Evidence that pricing algorithms collude comes from markets in which every seller runs the same algorithm. We ask what happens when they do not. Four language models from four providers, each at its cheapest tier, price in a four-firm logit Bertrand market without communication, in every homogeneous, two-by-two and fully mixed composi- tion (11 cells, 20 runs, 200 periods). Collusion is a property of the model: Claude and Gemini markets reach 72 to 79 percent of the monopoly rent, DeepSeek markets 24 percent, and GPT markets none, though GPT prices drift above the monopoly level rather than toward competition. Mixing does not reduce collusion by itself. Markets containing Gemini, which opens at the highest price and settles highest, are more col- lusive than the homogeneous markets they are built from; markets containing Claude, which opens lower and follows its rivals down, are less so; and only the fully mixed market is significantly less collusive than the average homogeneous one. Stability de- pends on the least stable participant: two GPT firms suffice to keep any market from converging. Inside mixed markets the rent is shared in a transitive order, DeepSeek over Claude over Gemini over GPT, that inverts the anchor ranking. The model that raises the price collects the least of the rent, as the price-leadership model predicts.

econ.GN↗