SearcharxivSearch

arXiv subjects

Michael K. Gilson

Publications and source records attributed to Michael K. Gilson.

8 recordsLinked to original sources

ToolMol: Evolutionary Agentic Framework for Multi-objective Drug Discovery

Advances in large language models (LLMs) have recently opened new and promising avenues for small-molecule drug discovery. Yet existing LLM-based approaches for molecular generation often suffer from high rates of invalid and low-quality ligand candidates, a result of the syntactic limitations of current models with regard to molecular strings. In this paper, we introduce $\texttt{ToolMol}$, an evolutionary agentic framework for de novo drug design. $\texttt{ToolMol}$ combines a multi-objective genetic algorithm with an agentic LLM operator that iteratively updates the ligand population. We build a comprehensive toolbox of RDKit-backed functions that allows our agentic operator to consisently make precise ligand modifications. $\texttt{ToolMol}$ achieves state-of-the-art performance on multi-objective property optimization tasks, discovering drug-like and synthesizable ligands that have $>10\%$ stronger predicted binding affinity compared to existing methods, evaluated on three protein targets. $\texttt{ToolMol}$ ligands additionally achieve state-of-the-art results in gold-standard Absolute Binding Free Energy scores, gaining over existing methods by over $35\%$. By studying chain-of-thought reasoning traces, we observe that tool-calling enables the model to more faithfully execute its planned modifications, efficiently exploiting the strong chemical prior knowledge in LLMs.

cs.LG

MF-LAL: Drug Compound Generation Using Multi-Fidelity Latent Space Active Learning

Current generative models for drug discovery primarily use molecular docking as an oracle to guide the generation of active compounds. However, such models are often not useful in practice because even compounds with high docking scores do not consistently show real-world experimental activity. More accurate methods for activity prediction exist, such as molecular dynamics based binding free energy calculations, but they are too computationally expensive to use in a generative model. To address this challenge, we propose Multi-Fidelity Latent space Active Learning (MF-LAL), a generative modeling framework that integrates a set of oracles with varying cost-accuracy tradeoffs. Using active learning, we train a surrogate model for each oracle and use these surrogates to guide generation of compounds with high predicted activity. Unlike previous approaches that separately learn the surrogate model and generative model, MF-LAL combines the generative and multi-fidelity surrogate models into a single framework, allowing for more accurate activity prediction and higher quality samples. Our experiments on two disease-relevant proteins show that MF-LAL produces compounds with significantly better binding free energy scores than other single and multi-fidelity approaches (~50% improvement in mean binding free energy score). The code is available at https://github.com/Rose-STL-Lab/MF-LAL.

cs.LG

Structure-Based Experimental Datasets for Benchmarking Protein Simulation Force Fields

This review article provides an overview of structurally oriented experimental datasets that can be used to benchmark protein force fields, focusing on data generated by nuclear magnetic resonance (NMR) spectroscopy and room temperature (RT) protein crystallography. We discuss what the observables are, what they tell us about structure and dynamics, what makes them useful for assessing force field accuracy, and how they can be connected to molecular dynamics simulations carried out using the force field one wishes to benchmark. We also touch on statistical issues that arise when comparing simulations with experiment. We hope this article will be particularly useful to computational researchers and trainees who develop, benchmark, or use protein force fields for molecular simulations.

q-bio.BM

Technical report: Improving the properties of molecules generated by LIMO

This technical report investigates variants of the Latent Inceptionism on Molecules (LIMO) framework to improve the properties of generated molecules. We conduct ablative studies of molecular representation, decoder model, and surrogate model training scheme. The experiments suggest that an autogressive Transformer decoder with GroupSELFIES achieves the best average properties for the random generation task.

cs.LG

Target-Free Compound Activity Prediction via Few-Shot Learning

Predicting the activities of compounds against protein-based or phenotypic assays using only a few known compounds and their activities is a common task in target-free drug discovery. Existing few-shot learning approaches are limited to predicting binary labels (active/inactive). However, in real-world drug discovery, degrees of compound activity are highly relevant. We study Few-Shot Compound Activity Prediction (FS-CAP) and design a novel neural architecture to meta-learn continuous compound activities across large bioactivity datasets. Our model aggregates encodings generated from the known compounds and their activities to capture assay information. We also introduce a separate encoder for the unknown compound. We show that FS-CAP surpasses traditional similarity-based techniques as well as other state of the art few-shot learning methods on a variety of target-free drug discovery settings and datasets.

cs.LG

LIMO: Latent Inceptionism for Targeted Molecule Generation

Generation of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily employ reinforcement learning, Markov sampling, or deep generative models guided by Gaussian processes, which can be prohibitively slow when generating molecules with high binding affinity calculated by computationally-expensive physics-based methods. We present Latent Inceptionism on Molecules (LIMO), which significantly accelerates molecule generation with an inceptionism-like technique. LIMO employs a variational autoencoder-generated latent space and property prediction by two neural networks in sequence to enable faster gradient-based reverse-optimization of molecular properties. Comprehensive experiments show that LIMO performs competitively on benchmark tasks and markedly outperforms state-of-the-art techniques on the novel task of generating drug-like compounds with high binding affinity, reaching nanomolar range against two protein targets. We corroborate these docking-based results with more accurate molecular dynamics-based calculations of absolute binding free energy and show that one of our generated drug-like compounds has a predicted $K_D$ (a measure of binding affinity) of $6 \cdot 10^{-14}$ M against the human estrogen receptor, well beyond the affinities of typical early-stage drug candidates and most FDA-approved drugs to their respective targets. Code is available at https://github.com/Rose-STL-Lab/LIMO.

cs.LG

Enhanced Diffusion and Chemotaxis of Enzymes

Many enzymes appear to diffuse faster in the presence of substrate and to drift either up or down a concentration gradient of their substrate. Observations of these phenomena, termed enhanced enzyme diffusion (EED) and enzyme chemotaxis, respectively, lead to a novel view of enzymes as active matter. Enzyme chemotaxis and EED may be important in biology, and they could have practical applications in biotechnology and nanotechnology. They also are of considerable biophysical interest; indeed, their physical mechanisms are still quite uncertain. This review provides an analytic summary of experimental studies of these phenomena and of the mechanisms that have been proposed to explain them, and offers a perspective of future directions for the field.

physics.chem-ph

Structure and Thermodynamics of Molecular Hydration via Grid Inhomogeneous Solvation Theory

Changes in hydration are central to the phenomenon of biomolecular recognition, but it has been difficult to properly frame and answer questions about their precise thermodynamic role. We address this problem by introducing Grid Inhomogeneous Solvation Theory (GIST), which discretizes the equations of Inhomogeneous Solvation Theory on a 3D grid in a volume of interest. Here, the solvent volume is divided into small grid boxes and localized thermodynamic entropies, energies and free energies are defined for each grid box. Thermodynamic solvation quantities are defined in such a manner that summing the quantities over all the grid boxes yields the desired total quantity for the system. This approach smoothly accounts for the thermodynamics of not only highly occupied water sites but also partly occupied and water depleted regions of the solvent, without the need for ad hoc terms drawn from other theories. The GIST method has the further advantage of allowing a rigorous end-states analysis that, for example in the problem of molecular recognition, can account for not only the thermodynamics of displacing water from the surface but also for the thermodynamics of solvent reorganization around the bound complex. As a preliminary application, we present GIST calculations at the 1-body level for the host cucurbit[7]uril, a low molecular weight receptor molecule which represents a tractable model for biomolecular recognition. One of the most striking results is the observation of a toroidal region of water density, at the center of the host's nonpolar cavity, which is significantly disfavored entropically, and hence may contribute to the ability of this small receptor to bind guest molecules with unusually high affinities.

q-bio.BM