SearcharxivSearch

arXiv subjects

Kresten Lindorff-Larsen

Publications and source records attributed to Kresten Lindorff-Larsen.

At least 19 recordsLinked to original sources

Detection of residual native state entropy changes upon mutation in Fyn SH3

NMR relaxation experiments have shown that there are small but measurable changes in the native state dynamics of the Fyn SH3 domain associated with the substitution by other amino acids of a phenylalanine residue (F20) in the hydrophobic core. We have here used experimental values of NMR order parameters for the wild type protein and two mutational variants (F20L and F20V) as restraints in molecular dynamics simulations. This approach is highly sensitive and provides an atomistic description of the subtle perturbations in native state fluctuations accompanying the mutations. The structural ensembles that we have determined using this method allow the changes in the native state entropy of the protein caused by each of the mutations to be estimated. These entropy changes correspond to free energy variations of several kcal/mol and therefore represent sizable contributions to the overall changes in stability that are associated with the amino acid mutations.

q-bio.BM

Zero-shot protein stability prediction by inverse folding models: a free energy interpretation

Inverse folding models have proven to be highly effective zero-shot predictors of protein stability. Despite this success, the link between the amino acid preferences of an inverse folding model and the free-energy considerations underlying thermodynamic stability remains incompletely understood. A better understanding would be of interest not only from a theoretical perspective, but also potentially provide the basis for stronger zero-shot stability prediction. In this paper, we take steps to clarify the free-energy foundations of inverse folding models. Our derivation reveals the standard practice of likelihood ratios as a simplistic approximation and suggests several paths towards better estimates of the relative stability. We empirically assess these approaches and demonstrate that considerable gains in zero-shot performance can be achieved with fairly simple means.

cs.LG

Integrative modelling of biomolecular dynamics

Much of our mechanistic understanding of the functions of biological macromolecules is based on static structural experiments, which can be modelled either as single structures or conformational ensembles. While these provide us with invaluable insights, they do not directly reveal that molecules are inherently dynamic. Advances in time-dependent and time-resolved experimental methods have made it possible to capture the dynamics of biomolecules at increasingly higher spatial and temporal resolutions. To complement these, computational models can represent the structural and dynamical behaviour of biomolecules at atomistic resolution and femtosecond timescale, and are therefore useful to interpret these experiments. Here, we review the progress in integrating simulations with dynamical experiments, focusing on the combination of simulations with time-resolved and time-dependent experimental data.

q-bio.BM

Computational design of intrinsically disordered proteins

Protein design has the potential to revolutionize biotechnology and medicine. While most efforts have focused on proteins with well-defined structures, increased recognition of the functional significance of intrinsically disordered regions, together with improvements in their modeling, has paved the way to their computational de novo design. This review summarizes recent advances in engineering intrinsically disordered regions with tailored conformational ensembles, molecular recognition, and phase behavior. We discuss challenges in combining models with predictive accuracy with scalable design workflows and outline emerging strategies that integrate knowledge-based, physics-based, and machine-learning approaches.

q-bio.BM

Towards a Unified Framework for Determining Conformational Ensembles of Disordered Proteins

Disordered proteins play essential roles in myriad cellular processes, yet their structural characterization remains a major challenge due to their dynamic and heterogeneous nature. We here present a community-driven initiative to address this problem by advocating a unified framework for determining conformational ensembles of disordered proteins. Our aim is to integrate state-of-the-art experimental techniques with advanced computational methods, including knowledge-based sampling, enhanced molecular dynamics, and machine learning models. The modular framework comprises three interconnected components: experimental data acquisition, computational ensemble generation, and validation. The systematic development of this framework will ensure the accurate and reproducible determination of conformational ensembles of disordered proteins. We highlight the open challenges necessary to achieve this goal, including force field accuracy, efficient sampling, and environmental dependency, advocating for collaborative benchmarking and standardized protocols.

q-bio.BM

How to use quantum computers for biomolecular free energies

Free energy calculations are at the heart of physics-based analyses of biochemical processes. They allow us to quantify molecular recognition mechanisms, which determine a wide range of biological phenomena from how cells send and receive signals to how pharmaceutical compounds can be used to treat diseases. Quantitative and predictive free energy calculations require computational models that accurately capture both the varied and intricate electronic interactions between molecules as well as the entropic contributions from motions of these molecules and their aqueous environment. However, accurate quantum-mechanical energies and forces can only be obtained for small atomistic models, not for large biomacromolecules. Here, we demonstrate how to consistently link accurate quantum-mechanical data obtained for substructures to the overall potential energy of biomolecular complexes by machine learning in an integrated algorithm. We do so using a two-fold quantum embedding strategy where the innermost quantum cores are treated at a very high level of accuracy. We demonstrate the viability of this approach for the molecular recognition of a ruthenium-based anticancer drug by its protein target, applying traditional quantum chemical methods. As such methods scale unfavorable with system size, we analyze requirements for quantum computers to provide highly accurate energies that impact the resulting free energies. Once the requirements are met, our computational pipeline FreeQuantum is able to make efficient use of the quantum computed energies, thereby enabling quantum computing enhanced modeling of biochemical processes. This approach combines the exponential speedups of quantum computers for simulating interacting electrons with modern classical simulation techniques that incorporate machine learning to model large molecules.

quant-ph

Software package for simulations using the coarse-grained CALVADOS model

We present the CALVADOS package for performing simulations of biomolecules using OpenMM and the coarse-grained CALVADOS model. The package makes it easy to run simulations using the family of CALVADOS models of biomolecules including disordered proteins, multi-domain proteins, proteins in crowded environments, and disordered RNA. We briefly describe the CALVADOS force fields and how they were parametrised. We then discuss the design paradigms and architecture of the CALVADOS package, and give examples of how to use it for running and analysing simulations. The simulation package is freely available under a GNU GPL license; therefore, it can easily be extended and we provide some examples of how this might be done.

q-bio.BM

The need to implement FAIR principles in biomolecular simulations

This letter illustrates the opinion of the molecular dynamics (MD) community on the need to adopt a new FAIR paradigm for the use of molecular simulations. It highlights the necessity of a collaborative effort to create, establish, and sustain a database that allows findability, accessibility, interoperability, and reusability of molecular dynamics simulation data. Such a development would democratize the field and significantly improve the impact of MD simulations on life science research. This will transform our working paradigm, pushing the field to a new frontier. We invite you to support our initiative at the MDDB community (https://mddbr.eu/community/) Now published as: Amaro, R.E., et al. The need to implement FAIR principles in biomolecular simulations. Nat Methods (2025) https://doi.org/10.1038/s41592-025-02635-0

q-bio.BM

Structure-Based Experimental Datasets for Benchmarking Protein Simulation Force Fields

This review article provides an overview of structurally oriented experimental datasets that can be used to benchmark protein force fields, focusing on data generated by nuclear magnetic resonance (NMR) spectroscopy and room temperature (RT) protein crystallography. We discuss what the observables are, what they tell us about structure and dynamics, what makes them useful for assessing force field accuracy, and how they can be connected to molecular dynamics simulations carried out using the force field one wishes to benchmark. We also touch on statistical issues that arise when comparing simulations with experiment. We hope this article will be particularly useful to computational researchers and trainees who develop, benchmark, or use protein force fields for molecular simulations.

q-bio.BM

Machine Learning Enhanced Calculation of Quantum-Classical Binding Free Energies

Binding free energies are a key element in understanding and predicting the strength of protein--drug interactions. While classical free energy simulations yield good results for many purely organic ligands, drugs including transition metal atoms often require quantum chemical methods for an accurate description. We propose a general and automated workflow that samples the potential energy surface with hybrid quantum mechanics/molecular mechanics (QM/MM) calculations and trains a machine learning (ML) potential on the QM energies and forces to enable efficient alchemical free energy simulations. To represent systems including many different chemical elements efficiently and to account for the different description of QM and MM atoms, we propose an extension of element-embracing atom-centered symmetry functions for QM/MM data as an ML descriptor. The ML potential approach takes electrostatic embedding and long-range electrostatics into account. We demonstrate the applicability of the workflow on the well-studied protein--ligand complex of myeloid cell leukemia 1 and the inhibitor 19G and on the anti-cancer drug NKP1339 acting on the glucose-regulated protein 78.

physics.chem-ph

Hierarchical quantum embedding by machine learning for large molecular assemblies

We present a quantum-in-quantum embedding strategy coupled to machine learning potentials to improve on the accuracy of quantum-classical hybrid models for the description of large molecules. In such hybrid models, relevant structural regions (such as those around reaction centers or pockets for binding of host molecules) can be described by a quantum model that is then embedded into a classical molecular-mechanics environment. However, this quantum region may become so large that only approximate electronic structure models are applicable. To then restore accuracy in the quantum description, we here introduce the concept of quantum cores within the quantum region that are amenable to accurate electronic structure models due to their limited size. Huzinaga-type projection-based embedding, for example, can deliver accurate electronic energies obtained with advanced electronic structure methods. The resulting total electronic energies are then fed into a transfer learning approach that efficiently exploits the higher-accuracy data to improve on a machine learning potential obtained for the original quantum-classical hybrid approach. We explore the potential of this approach in the context of a well-studied protein-ligand complex for which we calculate the free energy of binding using alchemical free energy and non-equilibrium switching simulations.

physics.chem-ph

Machine learning methods to study sequence-ensemble-function relationships in disordered proteins

Recent years have seen tremendous developments in the use of machine learning models to link amino acid sequence, structure and function of folded proteins. These methods are, however, rarely applicable to the wide range of proteins and sequences that comprise intrinsically disordered regions. We here review developments in the study of sequence-ensemble-function relationships of disordered proteins that exploit or are used to train machine learning models. These include methods for generating conformational ensembles and designing new sequences, and for linking sequences to biophysical properties and biological functions. We highlight how these developments are built on a tight integration between experiment, theory and simulations, and account for evolutionary constraints, which operate on sequences of disordered regions differently than on those of folded domains.

q-bio.BM

Guidelines for releasing a variant effect predictor

Computational methods for assessing the likely impacts of mutations, known as variant effect predictors (VEPs), are widely used in the assessment and interpretation of human genetic variation, as well as in other applications like protein engineering. Many different VEPs have been released to date, and there is tremendous variability in their underlying algorithms and outputs, and in the ways in which the methodologies and predictions are shared. This leads to considerable challenges for end users in knowing which VEPs to use and how to use them. Here, to address these issues, we provide guidelines and recommendations for the release of novel VEPs. Emphasising open-source availability, transparent methodologies, clear variant effect score interpretations, standardised scales, accessible predictions, and rigorous training data disclosure, we aim to improve the usability and interpretability of VEPs, and promote their integration into analysis and evaluation pipelines. We also provide a large, categorised list of currently available VEPs, aiming to facilitate the discovery and encourage the usage of novel methods within the scientific community.

q-bio.OT

Diffusion of intrinsically disordered proteins within viscoelastic membraneless droplets

In living cells, intrinsically disordered proteins (IDPs), such as FUS and DDX4, undergo phase separation, forming biomolecular condensates. Using molecular dynamics simulations, we investigate their behavior in their respective homogenous droplets. We find that the proteins exhibit transient subdiffusion due to the viscoelastic nature and confinement effects in the droplets. The conformation and the instantaneous diffusivity of the proteins significantly vary between the interior and the interface of the droplet, resulting in non-Gaussianity in the displacement distributions. This study highlights key aspects of IDP behavior in biomolecular condensates.

cond-mat.soft

Conformational ensembles of intrinsically disordered proteins and flexible multidomain proteins

Intrinsically disordered proteins (IDPs) and multidomain proteins with flexible linkers show a high level of structural heterogeneity and are best described by ensembles consisting of multiple conformations with associated thermodynamic weights. Determining conformational ensembles usually involves integration of biophysical experiments and computational models. In this review, we discuss current approaches to determining conformational ensembles of IDPs and multidomain proteins, including the choice of biophysical experiments, computational models used to sample protein conformations, models to calculate experimental observables from protein structure, and methods to refine ensembles against experimental data. We also provide examples of recent applications of integrative conformational ensemble determination to study IDPs and multidomain proteins and suggest future directions for research in the field.

q-bio.BM

On the potential of machine learning to examine the relationship between sequence, structure, dynamics and function of intrinsically disordered proteins

Intrinsically disordered proteins (IDPs) constitute a broad set of proteins with few uniting and many diverging properties. IDPs-and intrinsically disordered regions (IDRs) interspersed between folded domains-are generally characterized as having no persistent tertiary structure; instead they interconvert between a large number of different and often expanded structures. IDPs and IDRs are involved in an enormously wide range of biological functions and reveal novel mechanisms of interactions, and while they defy the common structure-function paradigm of folded proteins, their structural preferences and dynamics are important for their function. We here discuss open questions in the field of IDPs and IDRs, focusing on areas where machine learning and other computational methods play a role. We discuss computational methods aimed to predict transiently formed local and long-range structure, including methods for integrative structural biology. We discuss the many different ways in which IDPs and IDRs can bind to other molecules, both via short linear motifs, as well as in the formation of larger dynamic complexes such as biomolecular condensates. We discuss how experiments are providing insight into such complexes and may enable more accurate predictions. Finally, we discuss the role of IDPs in disease and how new methods are needed to interpret the mechanistic effects of genomic variants in IDPs.

q-bio.BM

Linking thermodynamics and measurements of protein stability

We review the background, theory and general equations for the analysis of equilibrium protein unfolding experiments, focusing on denaturant and heat-induced unfolding. The primary focus is on the thermodynamics of reversible folding/unfolding transitions and the experimental methods that are available for extracting thermodynamic parameters. We highlight the importance of modelling both how the folding equilibrium depends on a perturbing variable such as temperature or denaturant concentration, and the importance of modelling the baselines in the experimental observables.

physics.bio-ph

Dissecting the statistical properties of the Linear Extrapolation Method of determining protein stability

When protein stability is measured by denaturant induced unfolding the linear extrapolation method is usually used to analyse the data. This method is based on the observation that the change in Gibbs free energy associated with unfolding, $Δ_rG$, is often found to be a linear function of the denaturant concentration, $D$. The free energy change of unfolding in the absence of denaturant, $Δ_rG_0$, is estimated by extrapolation from this linear relationship. Data analysis is generally done by nonlinear least-squares regression to obtain estimates of the parameters as well as confidence intervals. We have compared different methods for calculating confidence intervals of the parameters and found that a simple method based on linear theory gives as good, if not better, results than more advanced methods. We have also compared three different parameterizations of the linear extrapolation method and show that one of the forms, $Δ_rG(D) = Δ_rG_0 - mD$, is problematic since the value of $Δ_rG_0$ and that of the $m$-value are correlated in the nonlinear least-squares analysis. Parameter correlation can in some cases cause problems in the estimation of confidence-intervals and -regions and should be avoided when possible. Two alternative parameterizations, $Δ_rG(D) = -m(D-D_{50})$ and $Δ_rG(D) = Δ_rG_0(1-D/D_{50})$, where $D_{50}$ is the midpoint of the transition region show much less correlation between parameters.

stat.ME