Searcharxiv⌕ Search

arXiv subjects

Peter C. St. John

Publications and source records attributed to Peter C. St. John.

4 recordsLinked to original sources

Plug & Play Directed Evolution of Proteins with Gradient-based Discrete MCMC

A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast MCMC sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

cs.LG↗

Message-passing neural networks for high-throughput polymer screening

Machine learning methods have shown promise in predicting molecular properties, and given sufficient training data machine learning approaches can enable rapid high-throughput virtual screening of large libraries of compounds. Graph-based neural network architectures have emerged in recent years as the most successful approach for predictions based on molecular structure, and have consistently achieved the best performance on benchmark quantum chemical datasets. However, these models have typically required optimized 3D structural information for the molecule to achieve the highest accuracy. These 3D geometries are costly to compute for high levels of theory, limiting the applicability and practicality of machine learning methods in high-throughput screening applications. In this study, we present a new database of candidate molecules for organic photovoltaic applications, comprising approximately 91,000 unique chemical structures.Compared to existing datasets, this dataset contains substantially larger molecules (up to 200 atoms) as well as extrapolated properties for long polymer chains. We show that message-passing neural networks trained with and without 3D structural information for these molecules achieve similar accuracy, comparable to state-of-the-art methods on existing benchmark datasets. These results therefore emphasize that for larger molecules with practical applications, near-optimal prediction results can be obtained without using optimized 3D geometry as an input. We further show that learned molecular representations can be leveraged to reduce the training data required to transfer predictions to a new DFT functional.

physics.comp-ph↗

Efficient estimation of the maximum metabolic productivity of batch systems

Production of chemicals from engineered organisms in a batch culture involves an inherent trade-off between productivity, yield, and titer. Existing strategies for strain design typically focus on designing mutations that achieve the highest yield possible while maintaining growth viability. While these methods are computationally tractable, an optimum productivity could be achieved by a dynamic strategy in which the intracellular division of resources is permitted to change with time. New methods for the design and implementation of dynamic microbial processes, both computational and experimental, have therefore been explored to maximize productivity. However, solving for the optimal metabolic behavior under the assumption that all fluxes in the cell are free to vary is a challenging numerical task. This work presents an efficient method for the calculation of a maximum theoretical productivity of a batch culture system using a dynamic optimization framework. This metric is analogous to the maximum theoretical yield, a measure that is well established in the metabolic engineering literature and whose use helps guide strain and pathway selection. The proposed method follows traditional assumptions of dynamic flux balance analysis: (1) that internal metabolite fluxes are governed by a pseudo-steady state, and (2) that external metabolite fluxes are dynamically bounded. The optimization is achieved via collocation on finite elements, and accounts explicitly for an arbitrary number of flux changes. The method can be further extended to explicitly solve for the trade-off curve between maximum productivity and yield. We demonstrate the method on succinate production in two common microbial hosts, Escherichia coli and Actinobacillus succinogenes, revealing that nearly optimal yields and productivities can be achieved with only two discrete flux stages.

q-bio.QM↗

A Coupled Stochastic Model Explains Differences in Circadian Behavior of Cry1 and Cry2 Knockouts

In the mammalian suprachiasmatic nucleus (SCN), a population of noisy cell-autonomous oscillators synchronizes to generate robust circadian rhythms at the organism-level. Within these cells two isoforms of Cryptochrome, Cry1 and Cry2, participate in a negative feedback loop driving circadian rhythmicity. Previous work has shown that single, dissociated SCN neurons respond differently to Cry1 and Cry2 knockouts: Cry1 knockouts are arrhythmic while Cry2 knockouts display more regular rhythms. These differences have led to speculation that CRY1 and CRY2 may play different functional roles in the oscillator. To address this proposition, we have developed a new coupled, stochastic model focused on the Period (Per) and Cry feedback loop, and incorporating intercellular coupling via vasoactive intestinal peptide (VIP). Due to the stochastic nature of molecular oscillations, we demonstrate that single-cell Cry1 knockout oscillations display partially rhythmic behavior, and cannot be classified as simply rhythmic or arrhythmic. Our model demonstrates that intrinsic molecular noise and differences in relative abundance, rather than differing functions, are sufficient to explain the range of rhythmicity encountered in Cry knockouts in the SCN. Our results further highlight the essential role of stochastic behavior in understanding and accurately modeling the circadian network and its response to perturbation.

q-bio.MN↗