SearcharxivSearch

arXiv subjects

G. Tiana

Publications and source records attributed to G. Tiana.

At least 19 recordsLinked to original sources

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

We investigate the parameter space of transformer models trained on protein sequence data using a statistical mechanics framework, sampling the loss landscape at varying temperatures by Langevin dynamics to characterize the low-loss manifold and understand the mechanisms underlying the superior performance of transformers in protein structure prediction. We find that, at variance with feedforward networks, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties. We show that the parameters of most layers are highly conserved at these temperatures if the dimension of the embedding is optimal, and we provide an operative way to find this dimension. Finally, we show that the attention matrix is more predictive of the contact maps of the protein at higher temperatures and for higher dimensions of the embedding than those optimal for learning.

cond-mat.dis-nn

The standard coil or globule phases cannot describe the denatured state of structured proteins and intrinsically disordered proteins

The concepts of globule and random coil were developed to describe the phases of homopolymers and then used to characterize the denatured state of structured cytosolic proteins and intrinsically disordered proteins. Using multi-scale molecular dynamics simulations, we were able to explore the conformational space of the disordered conformations of both types of protein under biological conditions in an affordable amount of computational time. By studying the size of the protein and the density correlations in space, we conclude that the standard phases of homopolymers and the tools to detect them cannot be applied straightforwardly to proteins.

q-bio.BM

Length-dependent residence time of contacts in simple polymeric models

Starting from the reported experimental evidence that the residence time of contacts between the ends of biopolymers is length dependent, we investigate the kinetics of contact breaking in simple polymer models from a theoretical point of view. We solved Kramers equation first for an ideal chain and then for a polymer with attracting ends, and compared the predictions with the results of molecular dynamics simulations. We found that the mean residence time always shows a power--law dependence on the length of the polymer with exponent $-1$, although is significantly smaller when obtained from the analysis of a single trajectory than when calculated from independent initial conformations. Only when the interaction is strong (>>kT) and the interaction range is small (of the order of the distance between consecutive monomers) does the residence time converge to that of the Arrhenius equation, independent of the length. We are able to provide expressions of the mean residence time for cases when the exact definition of contact is not available a priori, expressions that can be useful in typical cases of microscopy experiments.

cond-mat.soft

Predicting the binding of small molecules to proteins through invariant representation of the molecular structure

We present a computational scheme for predicting the ligands that bind to a pocket of known structure. It is based on the generation of a general abstract representation of the molecules, which is invariant to rotations, translations and permutations of atoms, and has some degree of isometry with the space of conformations. We use these representations to train a non-deep machine learning algorithm to classify the binding between pockets and molecule pairs, and show that this approach has a better generalization capability than existing methods.

physics.chem-ph

Structure of the space of folding protein sequences defined by large language models

Proteins populate a manifold in the high-dimensional sequence space whose geometrical structure guides their natural evolution. Leveraging recently-developed structure prediction tools based on transformer models, we first examine the protein sequence landscape as defined by the folding score function. This landscape shares characteristics with optimization challenges encountered in machine learning and constraint satisfaction problems. Our analysis reveals that natural proteins predominantly reside in wide, flat minima within this energy landscape. To investigate further, we employ statistical mechanics algorithms specifically designed to explore regions with high local entropy in relatively flat landscapes. Our findings indicate that these specialized algorithms can identify valleys with higher entropy compared to those found using traditional methods such as Monte Carlo Markov Chains. In a proof-of-concept case, we find that these highly entropic minima exhibit significant similarities to natural sequences, especially in critical key sites and local entropy. Additionally, evaluations through Molecular Dynamics suggests that the stability of these sequences closely resembles that of natural proteins. Our tool combines advancements in machine learning and statistical physics, providing new insights into the exploration of sequence landscapes where wide, flat minima coexist alongside a majority of narrower minima.

q-bio.BM

Locality of contacts determines the subdiffusion exponents in polymeric models of chromatin

Loop extrusion by motor proteins mediates the attractive interactions in chromatin on the length scale of megabases, providing the polymer with a well-defined structure and at the same time determining its dynamics. The mean square displacement of chromatin loci varies from a Rouse-like scaling to a more constrained subdiffusion, depending on cell type, genomic region and time scale. With a simple polymeric model, we show that such a Rouse-like dynamics occurs when the parameters of the model are chosen so that contacts are local along the chain, while in presence of non-local contacts, we observe subdiffusion at short time scales with exponents smaller than 0.5. Such exponents are independent of the detailed choice of the parameters and build a master curve that depends only on the mean locality of the resulting contacts. We compare the loop-extrusion model with a polymeric model with static links, showing that also in this case only the presence of nonlocal contacts can produce low-exponent subdiffusion. We interpret these results in terms of a simple analytical model.

q-bio.BM

Key interaction patterns in proteins revealed by cluster expansion of the partition function

The native conformation of structured proteins is stabilized by a complex network of interactions. We analyzed the elementary patterns that constitute such network and ranked them according to their importance in shaping protein sequence design. To achieve this goal, we employed a cluster expansion of the partition function in the space of sequences and evaluated numerically the statistical importance of each cluster. An important feature of this procedure is that it is applied to a dense, finite system. We found that patterns that contribute most to the partition function are cycles with even numbers of nodes, while cliques are typically detrimental. Each cluster also gives a contribute to the sequence entropy, which is a measure of the evolutionary designability of a fold. We compared the entropies associated with different interaction patterns to their abundances in the native structures of real proteins.

q-bio.BM

Effective Model of Loop Extrusion Predicts Chromosomal Domains

An active loop-extrusion mechanism is regarded as the main out--of--equilibrium mechanism responsible for the structuring of megabase-sized domains in chromosomes. We developed a model to study the dynamics of the chromosome fibre by solving the kinetic equations associated with the motion of the extruder. By averaging out the position of the extruder along the chain, we build an effective equilibrium model capable of reproducing experimental contact maps based solely on the positions of extrusion--blocking proteins. We assessed the quality of the effective model using numerical simulations of chromosomal segments and comparing the results with explicit-extruder models and experimental data.

q-bio.BM

Molecular recognition between cadherins studied by a coarse-grained model interacting with a coevolutionary potential

Studying the conformations involved in the dimerization of cadherins is highly relevant to understand the development of tissue and its failure, which is associated with tumors and metastases. Experimental techniques, like X-ray crystallography, can usually report only the most stable conformations, missing minority states that could nonetheless be important for the recognition mechanism. Computer simulations could be a valid complement to the experimental approach. However, standard all-atom protein models in explicit solvent are computationally too demanding to search thoroughly the conformational space of multiple chains composed of several hundreds of amino acids. To reach this goal, we resorted to a coarse-grained model in implicit solvent. The standard problem with this kind of models is to find a realistic potential to describe their interactions. We used coevolutionary information from cadherin alignments, corrected by a statistical potential, to build an interaction potential which is agnostic of the experimental conformations of the protein. Using this model, we explored the conformational space of multi-chain systems and validated the results comparing with experimental data. We identified dimeric conformations that are sequence-specific and that can be useful to rationalize the mechanism of recognition between cadherins.

q-bio.BM

Prediction of native contacts in proteins from an out--of--equilibrium coevolutionary process

The analysis of coevolution of residues in homologous proteins is a powerful tool to predict their native conformation. The standard framework in which coevolutionary analysis is usually worked out is that of equilibrium Potts models, assuming that proteins have evolved for enough time to reach thermodynamic equilibrium in sequence space. Here we propose a model to describe correlations in sequences based on an explicit description of the evolutionary kinetics of proteins. We show that this procedure improves the correct prediction of native contacts with respect to equilibrium--based models.

q-bio.BM

Statistical mechanical properties of sequence space determine the efficiency of the various algorithms to predict interaction energies and native contacts from protein coevolution

Studying evolutionary correlations in alignments of homologous sequences by means of an inverse Potts model has proven useful to obtain residue-residue contact energies and to predict contacts in proteins. The quality of the results depend much on several choices of the detailed model and on the algorithms used. We built, in a very controlled way, synthetic alignments with statistical properties similar to those of real proteins, and used them to assess the performance of different inversion algorithms and of their variants. Realistic synthetic alignments display typical features of low--temperature phases of disordered systems, a feature that affects the inversion algorithms. We showed that a Boltzmann--learning algorithm is computationally feasible and performs well in predicting the energy of native contacts. However, all algorithms suffer of false positives quite equally, making the quality of the prediction of native contacts with the different algorithm much system--dependent.

q-bio.BM

A method for partitioning the information contained in a protein sequence between its structure and function

Proteins employ the information stored in the genetic code and translated into their sequences to carry out well-defined functions in the cellular environment. The possibility to encode for such functions is controlled by the balance between the amount of information supplied by the sequence and that left after that the protein has folded into its structure. We developed a computational algorithm to evaluate the amount of information necessary to specify the protein structure, keeping into account the thermodynamic properties of protein folding. We thus show that the information remaining in the protein sequence after encoding for its structure (the 'information gap') is very close to what needed to encode for its function and interactions. Then, by predicting the information gap directly from the protein sequence, we show that it may be possible to use these insights from information theory to discriminate between ordered and disordered proteins, to identify unknown functions, and to optimize designed proteins sequences.

q-bio.BM

Properties of low-dimensional collective variables in the molecular dynamics of biopolymers

The description of the dynamics of a complex, high-dimensional system in terms of a low-dimensional set of collective variables Y can be fruitful if the low dimensional representation satisfies a Langevin equation with drift and diffusion coefficients which depend only on Y. We present a computational scheme to evaluate whether a given collective variable provides a faithful low-dimensional representation of the dynamics of a high-dimensional system. The scheme is based on the framework of finite-difference Langevin-equation, similar to that used for molecular-dynamics simulations. This allows one to calculate the drift and diffusion coefficients in any point of the full-dimensional system. The width of the distribution of drift and diffusion coefficients in an ensemble of microscopic points at the same value of Y indicates to which extent the dynamics of Y is described by a simple Langevin equation. Using a simple protein model we show that collective variables often used to describe biopolymers display a non-negligible width both in the drift and in the diffusion coefficients. We also show that the associated effective force is compatible with the equilibrium free--energy calculated from a microscopic sampling, but results in markedly different dynamical properties.

q-bio.BM

The looping probability of random heteropolymers helps to understand the scaling properties of biopolymers

Random heteropolymers are a minimal description of biopolymers and can provide a theoretical framework to the investigate the formation of loops in biophysical experiments. A two--state model provides a consistent and robust way to study the scaling properties of loop formation in polymers of the size of typical biological systems. Combining it with self--adjusting simulated--tempering simulations, we can calculate numerically the looping properties of several realizations of the random interactions within the chain. Differently from homopolymers, random heteropolymers display at different temperatures a continuous set of scaling exponents. The necessity of using self--averaging quantities makes finite--size effects dominant at low temperatures even for long polymers, shadowing the length--independent character of looping probability expected in analogy with homopolymeric globules. This could provide a simple explanation for the small scaling exponents found in experiments, for example in chromosome folding.

q-bio.QM

A many-body term improves the accuracy of effective potentials based on protein coevolutionary data

The study of correlated mutations in alignments of homologous proteins proved to be succesful not only in the prediction of their native conformation, but also in the developement of a two-body effective potential between pairs of amino acids. In the present work we extend the effective potential, introducing a many--body term based on the same theoretical framework, making use of a principle of maximum entropy. The extended potential performs better than the two--body one in predicting the energetic effect of 308 mutations in 14 proteins (including membrane proteins). The average value of the parameters of the many-body term correlates with the degree of hydrophobicity of the corresponding residues, suggesting that this term partly reflects the effect of the solvent.

q-bio.BM

Iterative derivation of effective potentials to sample the conformational space of proteins at atomistic scale

The current capacity of computers makes it possible to perform simulations of small systems with portable, explicit-solvent potentials achieving high degree of accuracy. However, simplified models must be employed to exploit the behaviour of large systems or to perform systematic scans of smaller systems. While powerful algorithms are available to facilitate the sampling of the conformational space, successful applications of such models are hindered by the availability of simple enough potentials able to satisfactorily reproduce known properties of the system. We develop an interatomic potential to account for a number of properties of proteins in a computationally economic way. The potential is defined within an all-atom, implicit solvent model by contact functions between the different atom types. The associated numerical values can be optimised by an iterative Monte Carlo scheme on any available experimental data, provided that they are expressible as thermal averages of some conformational properties. We test this model on three different proteins, for which we also perform a scan of all possible point mutations with explicit conformational sampling. The resulting models, optimised solely on a subset of native distances, not only reproduce the native conformations within a few Angstroms from the experimental ones, but show the cooperative transition between native and denatured state and correctly predict the measured free--energy changes associated with point mutations. Moreover, differently from other structure-based models, our method leaves a residual degree of frustration, which is known to be present in protein molecules.

q-bio.BM

Atomic-detailed milestones along the folding trajectory of protein G

The high computational cost of carrying out molecular dynamics simulations of even small-size proteins is a major obstacle in the study, at atomic detail and in explicit solvent, of the physical mechanism which is at the basis of the folding of proteins. Making use of a biasing algorithm, based on the principle of the ratchet-and-pawl, we have been able to calculate eight folding trajectories (to an RMSD between 1.2A and 2.5A) of the B1 domain of protein G in explicit solvent without the need of high-performance computing. The simulations show that in the denatured state there is a complex network of cause-effect relationships among contacts, which results in a rather hierarchical folding mechanism. The network displays few local and nonlocal native contacts which are cause of most of the others, in agreement with the NOE signals obtained in mildly-denatured conditions. Also nonnative contacts play an active role in the folding kinetics. The set of conformations corresponding to the transition state display phi-values with a correlation coefficient of 0.69 with the experimental ones. They are structurally quite homogeneous and topologically native-like, although some of the side chains and most of the hydrogen bonds are not in place.

q-bio.BM

Metadynamic sampling of the free energy landscapes of proteins coupled with a Monte Carlo algorithm

Metadynamics is a powerful computational tool to obtain the free energy landscape of complex systems. The Monte Carlo algorithm has proven useful to calculate thermodynamic quantities associated with simplified models of proteins, and thus to gain an ever-increasing understanding on the general principles underlying the mechanism of protein folding. We show that it is possible to couple metadynamics and Monte Carlo algorithms to obtain the free energy of model proteins in a way which is computationally very economical.

cond-mat.stat-mech