SearcharxivSearch

arXiv subjects

Justin Airas

Publications and source records attributed to Justin Airas.

2 recordsLinked to original sources

Knowledge Distillation of a Protein Language Model Yields a Foundational Implicit Solvent Model

Implicit solvent models (ISMs) promise to deliver the accuracy of explicit solvent simulations at a fraction of the computational cost. However, despite decades of development, their accuracy has remained insufficient for many critical applications, particularly for simulating protein folding and the behavior of intrinsically disordered proteins. Developing a transferable, data-driven ISM that overcomes the limitations of traditional analytical formulas remains a central challenge in computational chemistry. Here we address this challenge by introducing a novel strategy that distills the evolutionary information learned by a protein language model, ESM3, into a computationally efficient graph neural network (GNN). We show that this GNN potential, trained on effective energies from ESM3, is robust enough to drive stable, long-timescale molecular dynamics simulations. When combined with a standard electrostatics term, our hybrid model accurately reproduces protein folding free-energy landscapes and predicts the structural ensembles of intrinsically disordered proteins. This approach yields a single, unified model that is transferable across both folded and disordered protein states, resolving a long-standing limitation of conventional ISMs. By successfully distilling evolutionary knowledge into a physical potential, our work delivers a foundational implicit solvent model poised to accelerate the development of predictive, large-scale simulation tools.

physics.bio-ph

Scaling Graph Neural Networks to Large Proteins

Graph neural network (GNN) architectures have emerged as promising force field models, exhibiting high accuracy in predicting complex energies and forces based on atomic identities and Cartesian coordinates. To expand the applicability of GNNs, and machine learning force fields more broadly, optimizing their computational efficiency is critical, especially for large biomolecular systems in classical molecular dynamics simulations. In this study, we address key challenges in existing GNN benchmarks by introducing a dataset, DISPEF, which comprises large, biologically relevant proteins. DISPEF includes 207,454 proteins with sizes up to 12,499 atoms and features diverse chemical environments, spanning folded and disordered regions. The implicit solvation free energies, used as training targets, represent a particularly challenging case due to their many-body nature, providing a stringent test for evaluating the expressiveness of machine learning models. We benchmark the performance of seven GNNs on DISPEF, emphasizing the importance of directly accounting for long-range interactions to enhance model transferability. Additionally, we present a novel multiscale architecture, termed Schake, which delivers transferable and computationally efficient energy and force predictions for large proteins. Our findings offer valuable insights and tools for advancing GNNs in protein modeling applications.

physics.chem-ph