SearcharxivSearch

arXiv subjects

Michele Vendruscolo

Publications and source records attributed to Michele Vendruscolo.

At least 19 recordsLinked to original sources

Detection of residual native state entropy changes upon mutation in Fyn SH3

NMR relaxation experiments have shown that there are small but measurable changes in the native state dynamics of the Fyn SH3 domain associated with the substitution by other amino acids of a phenylalanine residue (F20) in the hydrophobic core. We have here used experimental values of NMR order parameters for the wild type protein and two mutational variants (F20L and F20V) as restraints in molecular dynamics simulations. This approach is highly sensitive and provides an atomistic description of the subtle perturbations in native state fluctuations accompanying the mutations. The structural ensembles that we have determined using this method allow the changes in the native state entropy of the protein caused by each of the mutations to be estimated. These entropy changes correspond to free energy variations of several kcal/mol and therefore represent sizable contributions to the overall changes in stability that are associated with the amino acid mutations.

q-bio.BM

ZeroFold: Protein-RNA Binding Affinity Predictions from Pre-Structural Embeddings

The accurate prediction of protein-RNA binding affinity remains an unsolved problem in structural biology, limiting opportunities in understanding gene regulation and designing RNA-targeting therapeutics. A central obstacle is the structural flexibility of RNA, as, unlike proteins, RNA molecules exist as dynamic conformational ensembles. Thus, committing to a single predicted structure discards information relevant to binding. Here, we show that this obstacle can be addressed by extracting pre-structural embeddings, which are intermediate representations from a biomolecular foundation model captured before the structure decoding step. Pre-structural embeddings implicitly encode conformational ensemble information without requiring predicted structures. We build ZeroFold, a transformer-based model that combines pre-structural embeddings from Boltz-2 for both protein and RNA molecules through a cross-modal attention mechanism to predict binding affinity directly from sequence. To support training and evaluation, we construct PRADB, a curated dataset of 2,621 unique protein-RNA pairs with experimentally measured affinities drawn from four complementary databases. On a held-out test set constructed with 40% sequence identity thresholds, ZeroFold achieves a Spearman correlation of 0.65, a value approaching the ceiling imposed by experimental measurement noise. Under progressively fairer evaluation conditions that control for training-set overlap, ZeroFold compares favourably with respect to leading structure-based and leading sequence-based predictors, with the performance gap widening as sequence similarity to competitor training data is reduced. These results illustrate how pre-structural embeddings offer a representation strategy for flexible biomolecules, opening a route to affinity prediction for protein-RNA pairs for which no structural data exist.

q-bio.BM

Multi-state Protein Design with DynamicMPNN

Structural biology has long been dominated by the one sequence, one structure, one function paradigm, yet many critical biological processes - from enzyme catalysis to membrane transport - depend on proteins that adopt multiple conformational states. Existing multi-state design approaches rely on post-hoc aggregation of single-state predictions, achieving poor experimental success rates compared to single-state design. We introduce DynamicMPNN, an inverse folding model explicitly trained to generate sequences compatible with multiple conformations through joint learning across conformational ensembles. Trained on 46,033 conformational pairs covering 75% of CATH superfamilies and evaluated using Alphafold 3, DynamicMPNN outperforms ProteinMPNN by up to 25% on decoy-normalized RMSD and by 12% on sequence recovery across our challenging multi-state protein benchmark.

cs.LG

Towards a Unified Framework for Determining Conformational Ensembles of Disordered Proteins

Disordered proteins play essential roles in myriad cellular processes, yet their structural characterization remains a major challenge due to their dynamic and heterogeneous nature. We here present a community-driven initiative to address this problem by advocating a unified framework for determining conformational ensembles of disordered proteins. Our aim is to integrate state-of-the-art experimental techniques with advanced computational methods, including knowledge-based sampling, enhanced molecular dynamics, and machine learning models. The modular framework comprises three interconnected components: experimental data acquisition, computational ensemble generation, and validation. The systematic development of this framework will ensure the accurate and reproducible determination of conformational ensembles of disordered proteins. We highlight the open challenges necessary to achieve this goal, including force field accuracy, efficient sampling, and environmental dependency, advocating for collaborative benchmarking and standardized protocols.

q-bio.BM

MD-LLM-1: A Large Language Model for Molecular Dynamics

Molecular dynamics (MD) is a powerful approach for modelling molecular systems, but it remains computationally intensive on spatial and time scales of many macromolecular systems of biological interest. To explore the opportunities offered by deep learning to address this problem, we introduce a Molecular Dynamics Large Language Model (MD-LLM) framework to illustrate how LLMs can be leveraged to learn protein dynamics and discover states not seen in training. By applying MD-LLM-1, the first implementation of this approach, obtained by fine-tuning Mistral 7B, to the T4 lysozyme and Mad2 protein systems, we show that training on one conformational state enables the prediction of other conformational states. These results indicate that MD-LLM-1 can learn the principles for the exploration of the conformational landscapes of proteins, although it is not yet modeling explicitly their thermodynamics and kinetics.

q-bio.BM

PLUMED Tutorials: a collaborative, community-driven learning ecosystem

In computational physics, chemistry, and biology, the implementation of new techniques in a shared and open source software lowers barriers to entry and promotes rapid scientific progress. However, effectively training new software users presents several challenges. Common methods like direct knowledge transfer and in-person workshops are limited in reach and comprehensiveness. Furthermore, while the COVID-19 pandemic highlighted the benefits of online training, traditional online tutorials can quickly become outdated and may not cover all the software's functionalities. To address these issues, here we introduce ``PLUMED Tutorials'', a collaborative model for developing, sharing, and updating online tutorials. This initiative utilizes repository management and continuous integration to ensure compatibility with software updates. Moreover, the tutorials are interconnected to form a structured learning path and are enriched with automatic annotations to provide broader context. This paper illustrates the development, features, and advantages of PLUMED Tutorials, aiming to foster an open community for creating and sharing educational resources.

physics.ed-ph

A method of incorporating rate constants as kinetic constraints in molecular dynamics simulations

From the point of view of statistical mechanics, a full characterisation of a molecular system requires the experimental determination of its possible states, their populations and the respective interconversion rates. Well-established methods can incorporate in molecular dynamics simulations experimental information about states using structural restraints, and about populations using thermodynamic restraints. However, it is still unclear how to include experimental knowledge of interconversion rates. Here we introduce a method of imposing known rate constants as constraints in molecular dynamics simulations, which is based on a combination of the maximum entropy and maximum caliber principles. Starting from an existing ensemble of trajectories, obtained from either molecular dynamics or enhanced trajectory sampling, this method provides a minimally perturbed path distribution consistent with the kinetic constraints, as well as a modified free energy and committor landscape. We illustrate the application of the method to simple toy systems, as well as to all atom molecular simulations of peptide association and folding. We find that by combining experimental rate coefficient data and molecular dynamics simulations we are able to determine new transition states, reaction mechanisms and free energies. For instance, in the case of chignolin protein folding we find that imposing a slower folding rate shifts the transition state to more native like conformations, while it increases the stability of the unfolded region.

physics.comp-ph

A practical guide to the simultaneous determination of protein structure and dynamics using metainference

Accurate protein structural ensembles can be determined with metainference, a Bayesian inference method that integrates experimental information with prior knowledge of the system and deals with all sources of uncertainty and errors as well as with system heterogeneity. Furthermore, metainference can be implemented using the metadynamics approach, which enables the computational study of complex biological systems requiring extensive conformational sampling. In this chapter, we provide a step-by-step guide to perform and analyse metadynamic metainference simulations using the ISDB module of the open-source PLUMED library, as well as a series of practical tips to avoid common mistakes. Specifically, we will guide the reader in the process of learning how to model the structural ensemble of a small disordered peptide by combining state-of-the-art molecular mechanics force fields with nuclear magnetic resonance data, including chemical shifts, scalar couplings and residual dipolar couplings.

q-bio.QM

Determination of protein structural ensembles using cryo-electron microscopy

Achieving a comprehensive understanding of the behaviour of proteins is greatly facilitated by the knowledge of their structures, thermodynamics and dynamics. All this information can be provided in an effective manner in terms of structural ensembles. A structural ensemble can be obtained by determining the structures, populations and interconversion rates for all the main states that a protein can occupy. To achieve this goal, integrative methods that combine experimental and computational approaches provide powerful tools. Here we focus on cryo-electron microscopy, which has become over recent years an invaluable resource to bridge the gap from order to disorder in structural biology. In this review, we provide a perspective of the current challenges and opportunities in determining protein structural ensembles using integrative approaches that can combine cryo-electron microscopy data with other available sources of information, along with an overview of the tools available to the community.

q-bio.QM

Metainference: A Bayesian Inference Method for Heterogeneous Systems

Modelling a complex system is almost invariably a challenging task. The incorporation of experimental observations can be used to improve the quality of a model, and thus to obtain better predictions about the behavior of the corresponding system. This approach, however, is affected by a variety of different errors, especially when a system populates simultaneously an ensemble of different states and experimental data are measured as averages over such states. To address this problem we present a Bayesian inference method, called metainference, that is able to deal with errors in experimental measurements as well as with experimental measurements averaged over multiple states. To achieve this goal, metainference models a finite sample of the distribution of models using a replica approach, in the spirit of the replica-averaging modelling based on the maximum entropy principle. To illustrate the method we present its application to a heterogeneous model system and to the determination of an ensemble of structures corresponding to the thermal fluctuations of a protein molecule. Metainference thus provides an approach to model complex systems with heterogeneous components and interconverting between different states by taking into account all possible sources of errors.

physics.comp-ph

Particle-based simulations of steady-state mass transport at high Péclet numbers

Conventional approaches for simulating steady-state distributions of particles under diffusive and advective transport at high Péclet numbers involve solving the diffusion and advection equations in at least two dimensions. Here, we present an alternative computational strategy by combining a particle-based rather than a field-based approach with the initialisation of particles in proportion to their flux. This method allows accurate prediction of the steady state and is applicable even at high Péclet numbers where traditional particle-based Monte-Carlo methods starting from randomly initialised particle distributions fail. We demonstrate that generating a flux of particles according to a predetermined density and velocity distribution at a single fixed time and initial location allows for accurate simulation of mass transport under flow. Specifically, upon initialisation in proportion to their flux, these particles are propagated individually and detected by summing up their Monte-Carlo trajectories in predefined detection regions. We demonstrate quantitative agreement of the predicted concentration profiles with the results of experiments performed with fluorescent particles in microfluidic channels under continuous flow. This approach is computationally advantageous and readily allows non-trivial initial distributions to be considered. In particular, this method is highly suitable for simulating advective and diffusive transport in microfluidic devices.

physics.flu-dyn

Analysis of the hierarchical structure of the B. subtilis transcriptional regulatory network

The transcriptional regulation of gene expression is orchestrated by complex networks of interacting genes. Increasing evidence indicates that these transcriptional regulatory networks (TRNs) in bacteria have an inherently hierarchical architecture, although the design principles and the specific advantages offered by this type of organization have not yet been fully elucidated. In this study, we focussed on the hierarchical structure of the TRN of the gram-positive bacterium Bacillus subtilis and performed a comparative analysis with the TRN of the gram-negative bacterium Escherichia coli. Using a graph-theoretic approach, we organized the transcription factors (TFs) and sigma-factors in the TRNs of B. subtilis and E. coli into three hierarchical levels (Top, Middle and Bottom) and studied several structural and functional properties across them. In addition to many similarities, we found also specific differences, explaining the majority of them with variations in the distribution of sigma-factors across the hierarchical levels in the two organisms. We then investigated the control of target metabolic genes by transcriptional regulators to characterize the differential regulation of three distinct metabolic subsystems (catabolism, anabolism and central energy metabolism). These results suggest that the hierarchical architecture that we observed in B. subtilis represents an effective organization of its TRN to achieve flexibility in the response to diverse stimuli.

q-bio.MN

Nucleated polymerisation in the presence of pre-formed seed filaments

We revisit the classical problem of nucleated polymerisation and derive a range of exact results describing polymerisation in systems intermediate between the well-known limiting cases of a reaction starting from purely soluble material and for a reaction where no new growth nuclei are formed.

physics.bio-ph

A Condensation-Ordering Mechanism in Nanoparticle-Catalyzed Peptide Aggregation

Nanoparticles introduced in living cells are capable of strongly promoting the aggregation of peptides and proteins. We use here molecular dynamics simulations to characterise in detail the process by which nanoparticle surfaces catalyse the self- assembly of peptides into fibrillar structures. The simulation of a system of hundreds of peptides over the millisecond timescale enables us to show that the mechanism of aggregation involves a first phase in which small structurally disordered oligomers assemble onto the nanoparticle and a second phase in which they evolve into highly ordered beta-sheets as their size increases.

q-bio.BM

Stochastic reconstruction of protein structures from effective connectivity profiles

We discuss a stochastic approach for reconstructing the native structures of proteins from the knowledge of the "effective connectivity", which is a one-dimensional structural profile constructed as a linear combination of the eigenvectors of the contact map of the target structure. The structural profile is used to bias a search of the conformational space towards the target structure in a Monte Carlo scheme operating on a C_alpha-chain of uniform, finite thickness. Structure information thus enters the folding dynamics via the effective connectivity, but the interaction is not restricted to pairs of amino acids that form native contacts, resulting in a free energy landscape which does not rely on the assumption of minimal frustration. Moreover, effective connectivity vectors can be predicted more readily from the amino acid sequence of proteins than the corresponding contact maps, thus suggesting that the stochastic protocol presented here could be effectively combined with other current methods for predicting native structures.

physics.bio-ph

Generic Mechanism of Emergence of Amyloid Protofilaments from Disordered Oligomeric aggregates

The presence of oligomeric aggregates, which is often observed during the process of amyloid formation, has recently attracted much attention since it has been associated with neurodegenerative conditions such as Alzheimer's and Parkinson's diseases. We provide a description of a sequence-indepedent mechanism by which polypeptide chains aggregate by forming metastable oligomeric intermediate states prior to converting into fibrillar structures. Our results illustrate how the formation of ordered arrays of hydrogen bonds drives the formation of beta-sheets within the disordered oligomeric aggregates that form early under the effect of hydrophobic forces. Initially individual beta-sheets form with random orientations, which subsequently tend to align into protofilaments as their lengths increases. Our results suggest that amyloid aggregation represents an example of the Ostwald step rule of first order phase transitions by showing that ordered cross-beta structures emerge preferentially from disordered compact dynamical intermediate assemblies.

q-bio.BM

Self-Templated Nucleation in Peptide and Protein aggregation

Peptides and proteins exhibit a common tendency to assemble into highly ordered fibrillar aggregates, whose formation proceeds in a nucleation-dependent manner that is often preceded by the formation of disordered oligomeric assemblies. This process has received much attention because disordered oligomeric aggregates have been associated with neurodegenerative disorders such as Alzheimer's and Parkinson's diseases. Here we describe a self-templated nucleation mechanism that determines the transition between the initial condensation of polypeptide chains into disordered assemblies and their reordering into fibrillar structures. The results that we present show that at the molecular level this transition is due to the ability of polypeptide chains to reorder within oligomers into fibrillar assemblies whose surfaces act as templates that stabilise the disordered assemblies.

q-bio.BM

Importance of Metastable States in the Free Energy Landscapes of Polypeptide Chains

We show that the interplay between excluded volume effects, hydrophobicity, and hydrogen bonding of a tube-like representation of a polypeptide chain gives rise to free energy landscapes that exhibit a small number of metastable minima corresponding to common structural motifs observed in proteins. The complexity of the landscape increases only moderately with the length of the chain. Analysis of the temperature dependence of these landscapes reveals that the stability of specific metastable states is maximal at a temperature close to the mid-point of folding. These mestastable states are therefore likely to be of particular significance in determining the generic tendency of proteins to aggregate into potentially pathogenic agents.

cond-mat.soft