SearcharxivSearch

arXiv subjects

George Church

Publications and source records attributed to George Church.

6 recordsLinked to original sources

Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity

The rapid adoption of generative artificial intelligence (GenAI) in the biosciences is transforming biotechnology, medicine, and synthetic biology. Yet this advancement is intrinsically linked to new vulnerabilities, as GenAI lowers the barrier to misuse and introduces novel biosecurity threats, such as generating synthetic viral proteins or toxins. These dual-use risks are often overlooked, as existing safety guardrails remain fragile and can be circumvented through deceptive prompts or jailbreak techniques. In this Perspective, we first outline the current state of GenAI in the biosciences and emerging threat vectors ranging from jailbreak attacks and privacy risks to the dual-use challenges posed by autonomous AI agents. We then examine urgent gaps in regulation and oversight, drawing on insights from 130 expert interviews across academia, government, industry, and policy. A large majority ($\approx 76$\%) expressed concern over AI misuse in biology, and 74\% called for the development of new governance frameworks. Finally, we explore technical pathways to mitigation, advocating a multi-layered approach to GenAI safety. These defenses include rigorous data filtering, alignment with ethical principles during development, and real-time monitoring to block harmful requests. Together, these strategies provide a blueprint for embedding security throughout the GenAI lifecycle. As GenAI becomes integrated into the biosciences, safeguarding this frontier requires an immediate commitment to both adaptive governance and secure-by-design technologies.

cs.CR

CodonMPNN for Organism Specific and Codon Optimal Inverse Folding

Generating protein sequences conditioned on protein structures is an impactful technique for protein engineering. When synthesizing engineered proteins, they are commonly translated into DNA and expressed in an organism such as yeast. One difficulty in this process is that the expression rates can be low due to suboptimal codon sequences for expressing a protein in a host organism. We propose CodonMPNN, which generates a codon sequence conditioned on a protein backbone structure and an organism label. If naturally occurring DNA sequences are close to codon optimality, CodonMPNN could learn to generate codon sequences with higher expression yields than heuristic codon choices for generated amino acid sequences. Experiments show that CodonMPNN retains the performance of previous inverse folding approaches and recovers wild-type codons more frequently than baselines. Furthermore, CodonMPNN has a higher likelihood of generating high-fitness codon sequences than low-fitness codon sequences for the same protein sequence. Code is available at https://github.com/HannesStark/CodonMPNN.

cs.LG

Large Language Models in Drug Discovery and Development: From Disease Mechanisms to Clinical Trials

The integration of Large Language Models (LLMs) into the drug discovery and development field marks a significant paradigm shift, offering novel methodologies for understanding disease mechanisms, facilitating drug discovery, and optimizing clinical trial processes. This review highlights the expanding role of LLMs in revolutionizing various stages of the drug development pipeline. We investigate how these advanced computational models can uncover target-disease linkage, interpret complex biomedical data, enhance drug molecule design, predict drug efficacy and safety profiles, and facilitate clinical trial processes. Our paper aims to provide a comprehensive overview for researchers and practitioners in computational biology, pharmacology, and AI4Science by offering insights into the potential transformative impact of LLMs on drug discovery and development.

q-bio.QM

The time is ripe to reverse engineer an entire nervous system: simulating behavior from neural interactions

Just like electrical engineers understand how microprocessors execute programs in terms of how transistor currents are affected by their inputs, neuroscientists want to understand behavior production in terms of how neuronal outputs are affected by their inputs and internal states. This dependency of neuronal outputs on inputs can be described by a state-dependent input-output (IO)-function. However, to reliably identify these IO-functions, we need to perturb each input and combinations of inputs while observing all the outputs. Here, we argue that such completeness is possible in C. elegans; a complete description that goes all the way from the activity of every neuron to predict behavior. The established and growing toolkit of optophysiology can non-invasively capture and control every neuron's activity and scale to countless experiments. The information from many such experiments can be pooled while capturing the inter-individual variability because neuronal identity and function are largely conserved across individuals. Just like electrical engineers use transistor IO-functions to simulate program execution, we argue that neuronal IO-functions could be used to simulate the impressive breadth of brain states and behaviors of C. elegans.

q-bio.NC

Grand Challenges for Global Brain Sciences

The next grand challenges for society and science are in the brain sciences. A collection of 60+ scientists from around the world, together with 10+ observers from national, private, and foundations, spent two days together discussing the top challenges that we could solve as a global community in the next decade. We eventually settled on three challenges, spanning anatomy, physiology, and medicine. Addressing all three challenges requires novel computational infrastructure. The group proposed the advent of The International Brain Station (TIBS), to address these challenges, and launch brain sciences to the next level of understanding.

q-bio.NC

New Dark Matter Detectors using DNA or RNA for Nanometer Tracking

Weakly Interacting Massive Particles (WIMPs) may constitute most of the matter in the Universe. The ability to detect the directionality of recoil nuclei will considerably facilitate detection of WIMPs. In this paper we propose a novel type of dark matter detector: detectors made of DNA or RNA could provide nanometer resolution for tracking, an energy threshold of 0.5 keV, and can operate at room temperature. When a WIMP from the Galactic Halo elastically scatters off of a nucleus in the detector, the recoiling nucleus then traverses hundreds of strings of single stranded nucleic acids (ssNA) with known base sequences and severs ssNA strands along its trajectory. The location of the break can be identified by amplifying and identifying the segments of cut ssNA using techniques well known to biologists. Thus the path of the recoiling nucleus can be tracked to nanometer accuracy. In one such detector concept, the transducers are nanometer-thick Au-foils of 1m x 1m, and the direction of recoiling nuclei is measured by "NA Tracking Chamber" consisting of ordered array of ssNA strands. Polymerase Chain Reaction (PCR) and ssNA sequencing are used to read-out the detector. The proposed detector is smaller and cheaper than other alternatives: 1 kg of gold and 0.1 to 4 kg of ssNA (depending on length and strand density), packed into 0.01m$^3$, can be used to study 10 GeV WIMPs. A variety of other detector target elements could be used in this detector to optimize for different WIMP masses and to identify WIMP properties. By leveraging advances in molecular biology, we aim to achieve about 1,000-fold better spatial resolution than in conventional WIMP detectors at reasonable cost.

astro-ph.IM