SearcharxivSearch

arXiv subjects

Frauke Gräter

Publications and source records attributed to Frauke Gräter.

8 recordsLinked to original sources

Predicting directional flexibility in proteins

Predicting protein dynamics is a long-standing problem in computational structural biology. Often, protein function critically depends on local directed motions, such as hinge movements, catalytic loop rearrangements and domain reorientations, which can be characterized by directional flexibility and correlated structural motions of the protein backbone. While Molecular Dynamics (MD) simulations provide an established but often prohibitively expensive approach, recent deep generative models aim to reduce this cost by directly predicting conformational ensembles, emulating MD. However, due to their large size and the need to generate several states until the derived dynamical properties converge, these models remain expensive. In this work, we propose BackFlip-2: a fast SE(3)-equivariant graph neural network trained to directly predict dynamical descriptors, such as directional backbone flexibility and pairwise dynamic correlations, from an equilibrium structure. In a series of experiments, we show that our model matches the accuracy of substantially larger ensemble generation models while being orders of magnitude faster, and demonstrate that the proposed equivariant architecture is especially well-suited for capturing anisotropic motions in proteins. BackFlip-2 model weights, training and inference code are available at https://github.com/graeter-group/backflip.

q-bio.BM

SITH: A Quantum-Chemical Framework for Predicting Bond Destabilization in Stretched Molecules

Mechanical forces can selectively destabilize chemical bonds of molecular systems, particularly in biological and synthetic polymers. While experimental and theoretical methods have advanced our understanding of mechanochemical processes, predicting where energy concentrates within a molecule remains a significant challenge. To address this, we introduce SITH (Splitting Intramolecular Tension due to stretcHing), a novel method that decomposes the total electronic energy change of a stretched molecule into contributions from individual degrees of freedom -- such as bond lengths, angles, and dihedrals -- using numerical integration of the work-energy theorem. Unlike previous approaches that rely on harmonic approximations, SITH provides high accuracy and robustness to study the distribution of energies of stretched molecules up to a first bond cleavage. Although SITH uses 3N-6 degrees of freedom for the energy decomposition, we show that it can work even for ring structures like prolines. We apply SITH to a dataset of tripeptides and demonstrate that glycine and proline exhibit significantly different energy distributions in their Ca-C backbone bonds under tension: glycine stores more energy, making it more prone to rupture, while proline has the opposite behaviour. These findings reveal intrinsic differences in mechanochemical susceptibility across amino acids, offering more accurate predictions of bond rupture in proteins, and similarly in other (bio)polymers. SITH thus provides a powerful, interpretable tool for understanding energy distribution at the quantum level, with possible implications in mechanochemistry and force field validation.

physics.chem-ph

Learning Potential Energy Surfaces of Hydrogen Atom Transfer Reactions in Peptides

Hydrogen atom transfer (HAT) reactions are essential in many biological processes, such as radical migration in damaged proteins, but their mechanistic pathways remain incompletely understood. Simulating HAT is challenging due to the need for quantum chemical accuracy at biologically relevant scales; thus, neither classical force fields nor DFT-based molecular dynamics are applicable. Machine-learned potentials offer an alternative, able to learn potential energy surfaces (PESs) with near-quantum accuracy. However, training these models to generalize across diverse HAT configurations, especially at radical positions in proteins, requires tailored data generation and careful model selection. Here, we systematically generate HAT configurations in peptides to build large datasets using semiempirical methods and DFT. We benchmark three graph neural network architectures (SchNet, Allegro, and MACE) on their ability to learn HAT PESs and indirectly predict reaction barriers from energy predictions. MACE consistently outperforms the others in energy, force, and barrier prediction, achieving a mean absolute error of 1.13 kcal/mol on out-of-distribution DFT barrier predictions. Using molecular dynamics, we show our MACE potential is stable, reactive, and generalizes beyond training data to model HAT barriers in collagen I. This accuracy enables integration of ML potentials into large-scale collagen simulations to compute reaction rates from predicted barriers, advancing mechanistic understanding of HAT and radical migration in peptides. We analyze scaling laws, model transferability, and cost-performance trade-offs, and outline strategies for improvement by combining ML potentials with transition state search algorithms and active learning. Our approach is generalizable to other biomolecular systems, enabling quantum-accurate simulations of chemical reactivity in complex environments.

cs.LG

Learning conformational ensembles of proteins based on backbone geometry

Deep generative models have recently been proposed for sampling protein conformations from the Boltzmann distribution, as an alternative to often prohibitively expensive Molecular Dynamics simulations. However, current state-of-the-art approaches rely on fine-tuning pre-trained folding models and evolutionary sequence information, limiting their applicability and efficiency, and introducing potential biases. In this work, we propose a flow matching model for sampling protein conformations based solely on backbone geometry - BBFlow. We introduce a geometric encoding of the backbone equilibrium structure as input and propose to condition not only the flow but also the prior distribution on the respective equilibrium structure, eliminating the need for evolutionary information. The resulting model is orders of magnitudes faster than current state-of-the-art approaches at comparable accuracy, is transferable to multi-chain proteins, and can be trained from scratch in a few GPU days. In our experiments, we demonstrate that the proposed model achieves competitive performance with reduced inference time, across not only an established benchmark of naturally occurring proteins but also de novo proteins, for which evolutionary information is scarce or absent. BBFlow is available at https://github.com/graeter-group/bbflow.

q-bio.BM

Flexibility-Conditioned Protein Structure Design with Flow Matching

Recent advances in geometric deep learning and generative modeling have enabled the design of novel proteins with a wide range of desired properties. However, current state-of-the-art approaches are typically restricted to generating proteins with only static target properties, such as motifs and symmetries. In this work, we take a step towards overcoming this limitation by proposing a framework to condition structure generation on flexibility, which is crucial for key functionalities such as catalysis or molecular recognition. We first introduce BackFlip, an equivariant neural network for predicting per-residue flexibility from an input backbone structure. Relying on BackFlip, we propose FliPS, an SE(3)-equivariant conditional flow matching model that solves the inverse problem, that is, generating backbones that display a target flexibility profile. In our experiments, we show that FliPS is able to generate novel and diverse protein backbones with the desired flexibility, verified by Molecular Dynamics (MD) simulations. FliPS and BackFlip are available at https://github.com/graeter-group/flips .

q-bio.BM

Generating Highly Designable Proteins with Geometric Algebra Flow Matching

We introduce a generative model for protein backbone design utilizing geometric products and higher order message passing. In particular, we propose Clifford Frame Attention (CFA), an extension of the invariant point attention (IPA) architecture from AlphaFold2, in which the backbone residue frames and geometric features are represented in the projective geometric algebra. This enables to construct geometrically expressive messages between residues, including higher order terms, using the bilinear operations of the algebra. We evaluate our architecture by incorporating it into the framework of FrameFlow, a state-of-the-art flow matching model for protein backbone generation. The proposed model achieves high designability, diversity and novelty, while also sampling protein backbones that follow the statistical distribution of secondary structure elements found in naturally occurring proteins, a property so far only insufficiently achieved by many state-of-the-art generative models.

cs.LG

Grappa -- A Machine Learned Molecular Mechanics Force Field

Simulating large molecular systems over long timescales requires force fields that are both accurate and efficient. In recent years, E(3) equivariant neural networks have lifted the tension between computational efficiency and accuracy of force fields, but they are still several orders of magnitude more expensive than established molecular mechanics (MM) force fields. Here, we propose Grappa, a machine learning framework to predict MM parameters from the molecular graph, employing a graph attentional neural network and a transformer with symmetry-preserving positional encoding. The resulting Grappa force field outperformstabulated and machine-learned MM force fields in terms of accuracy at the same computational efficiency and can be used in existing Molecular Dynamics (MD) engines like GROMACS and OpenMM. It predicts energies and forces of small molecules, peptides, RNA and - showcasing its extensibility to uncharted regions of chemical space - radicals at state-of-the-art MM accuracy. We demonstrate Grappa's transferability to macromolecules in MD simulations from a small fast folding protein up to a whole virus particle. Our force field sets the stage for biomolecular simulations closer to chemical accuracy, but with the same computational cost as established protein force fields.

physics.chem-ph

Mechanoradicals in tensed tendon collagen as a new source of oxidative stress

As established nearly a century ago, mechanoradicals originate from homolytic bond scission in polymers. The existence, nature and biological relevance of mechanoradicals in proteins, instead, are unknown. We here show that mechanical stress on collagen produces radicals and subsequently reactive oxygen species, essential biological signaling molecules. Electron-paramagnetic resonance (EPR) spectroscopy of stretched rat tail tendon, atomistic Molecular Dynamics simulations and quantum calculations show that the radicals form by bond scission in the direct vicinity of crosslinks in collagen. Radicals migrate to adjacent clusters of aromatic residues and stabilize on oxidized tyrosyl radicals, giving rise to a distinct EPR spectrum consistent with a stable dihydroxyphenylalanine (DOPA) radical. The protein mechanoradicals, as a yet undiscovered source of oxidative stress, finally convert into hydrogen peroxide. Our study suggests collagen I to have evolved as a radical sponge against mechano-oxidative damage and proposes a new mechanism for exercise-induced oxidative stress and redox-mediated pathophysiological processes.

physics.bio-ph