SearcharxivSearch

arXiv subjects

Guido Tiana

Publications and source records attributed to Guido Tiana.

16 recordsLinked to original sources

The energy landscape of DNA-binding proteins along the genome

Reconstructing the energy profile of DNA-binding proteins along the genome requires an algorithm that quantifies efficiently the binding free energy. We assembled a dataset of protein structures and DNA binding sites, together with their binding energies, and used it to train a machine-learning algorithm that learns a latent invariant representation of the protein interface and of the DNA sequence, combining them to predict the free energy. After validating the method, we used it to determine the energy profile of a single-domain transcription factor (PU.1) sliding on mammalian chromosomes, predicting its binding regions and quantifying the statistical properties that determine their stability and their kinetic accessibility.

q-bio.BM

Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches

Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for realistic datasets because they require repeated evaluation of full-batch gradients. We introduce a pseudo-Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using minibatches in a controlled manner. The method exploits the statistical properties of minibatch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains high computational diffusion and scales favorably to networks with over one million parameters. Finally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to SGD, without requiring a validation set or early stopping procedure. These results establish controlled minibatch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.

cond-mat.dis-nn

Sampling the space of solutions of an artificial neural network

The weight space of an artificial neural network can be systematically explored using tools from statistical mechanics. We employ a combination of a hybrid Monte Carlo algorithm which performs long exploration steps, a ratchet-based algorithm to investigate connectivity paths, and coupled replica models simulations to study subdominant flat regions. Our analysis focuses on one hidden layer networks and spans a range of energy levels and constrained density regimes. Near the interpolation threshold, the low-energy manifold shows a spiky topology. In the overparameterized regime, however, the low-energy manifold becomes entirely flat, forming an extended complex structure that is easy to sample. These numerical results are supported by an analytical study of the training error landscape, and we show numerically that the qualitative features of the loss landscape are robust across different data structures. Our study aims to provide new methodological insights for developing scalable methods for large networks.

cond-mat.dis-nn

Effective model of protein--mediated interactions in chromatin

Protein-mediated interactions are ubiquitous in the cellular environment, and particularly in the nucleus, where they are responsible for the structuring of chromatin. We show through molecular--dynamics simulations of a polymer surrounded by binders that the strength of the binder-polymer interaction separates an equilibrium from a non-equilibrium regime. In the equilibrium regime, the system can be efficiently described by an effective model in which the binders are traced out. Even in this case, the polymer display features that are different from those of a standard homopolymer interacting with two-body interactions. We then extend the effective model to deal with the case where binders cannot be regarded as in equilibrium and a new phenomenology appears, including local blobs in the polymer. Providing an effective description of the system can be useful in clarifying the fundamental mechanisms governing chromatin structuring.

q-bio.BM

Native state of natural proteins optimises local entropy

The differing ability of polypeptide conformations to act as the native state of proteins has long been rationalized in terms of differing kinetic accessibility or thermodynamic stability. Building on the successful applications of physical concepts and sampling algorithms recently introduced in the study of disordered systems, in particular artificial neural networks, we quantitatively explore how well a quantity known as the local entropy describes the native state of model proteins. In lattice models and all-atom representations of proteins, we are able to efficiently sample high local entropy states and to provide a proof of concept of enhanced stability and folding rate. Our methods are based on simple and general statistical--mechanics arguments, and thus we expect that they are of very general use.

q-bio.BM

Bifractal nature of chromosome contact maps

Modern biological techniques such as Hi-C permit to measure probabilities that different chromosomal regions are close in space. These probabilities can be visualised as matrices called contact maps. In this paper, we introduce a multifractal analysis of chromosomal contact maps. Our analysis reveals that Hi-C maps are bifractal, i.e. complex geometrical objects characterized by two distinct fractal dimensions. To rationalize this observation, we introduce a model that describes chromosomes as a hierarchical set of nested domains and we solve it exactly. The predicted multifractal spectrum is in excellent quantitative agreement with experimental data. Moreover, we show that our theory yields to a more robust estimation of the scaling exponent of the contact probability than existing methods. By applying this method to experimental data, we detect subtle conformational changes among chromosomes during differentiation of human stem cells.

q-bio.BM

Assessing the accuracy of direct-coupling analysis for RNA contact prediction

Many non-coding RNAs are known to play a role in the cell directly linked to their structure. Structure prediction based on the sole sequence is however a challenging task. On the other hand, thanks to the low cost of sequencing technologies, a very large number of homologous sequences are becoming available for many RNA families. In the protein community, it has emerged in the last decade the idea of exploiting the covariance of mutations within a family to predict the protein structure using the direct-coupling-analysis (DCA) method. The application of DCA to RNA systems has been limited so far. We here perform an assessment of the DCA method on 17 riboswitch families, comparing it with the commonly used mutual information analysis and with state-of-the-art R-scape covariance method. We also compare different flavors of DCA, including mean-field, pseudo-likelihood, and a proposed stochastic procedure (Boltzmann learning) for solving exactly the DCA inverse problem. Boltzmann learning outperforms the other methods in predicting contacts observed in high resolution crystal structures.

q-bio.QM

An implementation of the maximum-caliber principle by replica-averaged time-resolved restrained simulations

Inferential methods can be used to integrate experimental informations and molecular simulations. The maximum entropy principle provides a framework for using equilibrium experimental data and it has been shown that replica-averaged simulations, restrained using a static potential, are a practical and powerful implementation of such principle. Here we show that replica-averaged simulations restrained using a time-dependent potential are equivalent to the principle of maximum caliber, the dynamic version of the principle of maximum entropy, and thus may allow to integrate time-resolved data in molecular dynamics simulations. We provide an analytical proof of the equivalence as well as a computational validation making use of simple models and synthetic data. Some limitations and possible solutions are also discussed.

q-bio.BM

Spontaneous domain formation in disordered copolymers as a mechanism for chromosome structuring

Motivated by the problem of domain formation in chromosomes, we studied a co--polymer model where only a subset of the monomers feel attractive interactions. These monomers are displaced randomly from a regularly-spaced pattern, thus introducing some quenched disorder in the system. Previous work has shown that in the case of regularly-spaced interacting monomers this chain can fold into structures characterized by multiple distinct domains of consecutive segments. In each domain, attractive interactions are balanced by the entropy cost of forming loops. We show by advanced replica-exchange simulations that adding disorder in the position of the interacting monomers further stabilizes these domains. The model suggests that the partitioning of the chain into well-defined domains of consecutive monomers is a spontaneous property of heteropolymers. In the case of chromosomes, evolution could have acted on the spacing of interacting monomers to modulate in a simple way the underlying domains for functional reasons.

q-bio.BM

Complete coverage of space favors modularity of the grid system in the brain

Grid cells in the entorhinal cortex fire when animals that are exploring a certain region of space occupy the vertices of a triangular grid that spans the environment. Different neurons feature triangular grids that differ in their properties of periodicity, orientation and ellipticity. Taken together, these grids allow the animal to maintain an internal, mental representation of physical space. Experiments show that grid cells are modular, i.e. there are groups of neurons which have grids with similar periodicity, orientation and ellipticity. We use statistical physics methods to derive a relation between variability of the properties of the grids within a module and the range of space that can be covered completely (i.e. without gaps) by the grid system with high probability. Larger variability shrinks the range of representation, providing a functional rationale for the experimentally observed co-modularity of grid cell periodicity, orientation and ellipticity. We obtain a scaling relation between the number of neurons and the period of a module, given the variability and coverage range. Specifically, we predict how many more neurons are required at smaller grid scales than at larger ones.

q-bio.NC

The effect of disorder in the contact probability of elongated conformations of biopolymers

Biopolymers are characterized by heterogeneous interactions, and usually perform their biological tasks forming contacts within domains of limited size. Combining polymer theory with a replica approach, we study the scaling properties of the probability of contact formation in random heteropolymers as a function of their linear distance. It is found that close or above the theta--point, it is possible to define a contact probability which is typical (i.e. "self-averaging") for different realizations of the heterogeneous interactions, and which displays an exponential cut--off, dependent on temperature and on the interaction range. In many cases this cut--off is comparable with the typical sizes of domains in biopolymers. While it is well known that disorder causes interesting effects at low temperature, the behavior elucidated in the present study is an example of a non--trivial effect at high temperature.

q-bio.BM

The network of stabilizing contacts in proteins studied by coevolutionary data

The primary structure of proteins, that is their sequence, represents one of the most abundant set of experimental data concerning biomolecules. The study of correlations in families of co--evolving proteins by means of an inverse Ising--model approach allows to obtain information on their native conformation. Following up on a recent development along this line, we optimize the algorithm to calculate effective energies between the residues, validating the approach both back-calculating interaction energies in a model system, and predicting the free energies associated to mutations in real systems. Making use of these effective energies, we study the networks of interactions which stabilizes the native conformation of some well--studied proteins, showing that it display different properties than the associated contact network.

q-bio.BM

Ratcheted molecular-dynamics simulations identify efficiently the transition state of protein folding

The atomistic characterization of the transition state is a fundamental step to improve the understanding of the folding mechanism and the function of proteins. From a computational point of view, the identification of the conformations that build out the transition state is particularly cumbersome, mainly because of the large computational cost of generating a statistically-sound set of folding trajectories. Here we show that a biasing algorithm, based on the physics of the ratchet-and-pawl, can be used to identify efficiently the transition state. The basic idea is that the algorithmic ratchet exerts a force on the protein when it is climbing the free-energy barrier, while it is inactive when it is descending. The transition state can be identified as the point of the trajectory where the ratchet changes regime. Besides discussing this strategy in general terms, we test it within a protein model whose transition state can be studied independently by plain molecular dynamics simulations. Finally, we show its power in explicit-solvent simulations, obtaining and characterizing a set of transition--state conformations for ACBP and CI2.

q-bio.BM

Equilibrium properties of realistic random heteropolymers and their relevance for globular and naturally unfolded proteins

Random heteropolymers do not display the typical equilibrium properties of globular proteins, but are the starting point to understand the physics of proteins and, in particular, to describe their non-native states. So far, they have been studied only with mean-field models in the thermodynamic limit, or with computer simulations of very small chains on lattice. After describing a self-adjusting parallel-tempering technique to sample efficiently the low-energy states of frustrated systems without the need of tuning the system-dependent parameters of the algorithm, we apply it to random heteropolymers moving in continuous space. We show that if the mean interaction between monomers is negative, the usual description through the random energy model is nearly correct, provided that it is extended to account for non-compact conformations. If the mean interaction is positive, such a simple description breaks out and the system behaves in a way more similar to Ising spin glasses. The former case is a model for the denatured state of glob- ular proteins, the latter of naturally-unfolded proteins, whose equilibrium properties thus result qualitatively different.

q-bio.BM

Estimation of microscopic averages from metadynamics

With the help of metadynamics it is possible to calculate efficiently the free energy of systems displaying high energy barriers as a function of few selected "collective variables". In doing this, the contribution of all the other degrees of freedom ("microscopic" variables) is averaged out and, thus, lost. In the following, it is shown that it is possible to calculate the thermal average of these microscopic degrees of freedom during the metadynamics, not loosing this piece of information.

cond-mat.stat-mech

Hierarchy of events in protein folding: beyond the Go model

Simplified Go models, where only native contacts interact favorably, have proven useful to characterize some aspects of the folding of small proteins. The success of these models is limited by the fact that all residues interact in the same way, so that the folding features of a protein are determined only by the geometry of its native conformation. We present an extended version of a C-alpha based Go model where different residues interact with different energies. The model is used to calculate the thermodynamics of three small proteins (Protein G, SrcSH3 and CI2) and the effect of mutations on the wildtype sequence. The model allows to investigate some of the most controversial areas in protein folding such as its earliest stages, a subject which has lately received particular attention. The picture which emerges for the three proteins under study is that of a hierarchical process, where local elementary structures (LES) (not necessarily coincident with elements of secondary structure) are formed at the early stages of the folding and drive the protein, through the transition state and the postcritical folding nucleus (FN), resulting from the docking of the LES, to the native conformation.

q-bio.BM