Searcharxiv⌕ Search

arXiv subjects

Andrea Giansanti

Publications and source records attributed to Andrea Giansanti.

At least 19 recordsLinked to original sources

Insights into the role of dynamical features in protein complex formation: the case of SARS-CoV-2 spike binding with ACE2

The functionality of protein-protein complexes is closely tied to the strength of their interactions, making the evaluation of binding affinity a central focus in structural biology. However, the molecular determinants underlying binding affinity are still not fully understood. In particular, the entropic contributions, especially those arising from conformational dynamics, remain poorly characterized. In this study, we explore the relationship between protein motion and binding stability and its role in protein function. To gain deeper insight into how protein complexes modulate their stability, we investigated a model system with a well-characterized and fast evolutionary history: a set of SARS-CoV-2 spike protein variants bound to the human ACE2 receptor, for which experimental binding affinity data are available. Through Molecular Dynamics simulations, we analyzed both structural and dynamical differences between the unbound (apo) and bound (holo) forms of the spike protein across several variants of concern. Our findings indicate that a more stable binding is associated with proteins that exhibit higher rigidity in their unbound state and display dynamical patterns similar to that observed after binding to ACE2. The increase of binding stability is not the sole driving force of SARS-CoV-2 evolution. More recent variants are characterized by a more dynamical behavior that determines a less efficient viral entry but could optimize other traits, such as antibody escape. These results suggest that to fully understand the strength of the binding between two proteins, the stability of the two isolated partners should be investigated.

q-bio.BM↗

SOME Special Solutions of a Nonlinear System of 4 Ordinary Differential Equations Recently Introduced to Investigate the Evolution of Human Respiratory Virus Epidemics

A system of 4 nonlinearly-coupled Ordinary Differential Equations has been recently introduced to investigate the evolution of human respiratory virus epidemics. In this paper we point out that some explicit solutions of that system can be obtained by algebraic operations, provided the parameters of the model satisfy certain constraints.

math.CA↗

Fluctuation-response theorem for Kullback-Leibler divergences to quantify causation

We define a new measure of causation from a fluctuation-response theorem for Kullback-Leibler divergences, based on the information-theoretic cost of perturbations. This information response has both the invariance properties required for an information-theoretic measure and the physical interpretation of a propagation of perturbations. In linear systems, the information response reduces to the transfer entropy, providing a connection between Fisher and mutual information.

cs.IT↗

Co-evolution between Codon Usage and Protein-Protein Interaction in Bacteria

We study the correlation between the codon usage bias of genetic sequences and the network features of protein-protein interaction (PPI) in bacterial species. We use PCA techniques in the space of codon bias indices to show that genes with similar patterns of codon usage have a significantly higher probability that their encoded proteins are functionally connected and interacting. Importantly, this signal emerges when multiple aspects of codon bias are taken into account at the same time. The present study extends our previous observations on E.Coli over a wide set of 34 bacteria. These findings could allow for future investigations on the possible effects of codon bias on the topology of the PPI network, with the aim of improving existing bioinformatics methods for predicting protein interactions.

q-bio.GN↗

Bacterial protein interaction networks: connectivity is ruled by gene conservation, essentiality and function

Protein-protein interaction (PPI) networks are the backbone of all processes in living cells. In this work we relate conservation, essentiality and functional repertoire of a gene to the connectivity $k$ (i.e., the number of interaction links) of the corresponding protein in the PPI network. On a set of 42 bacterial genomes of different sizes, and with reasonably separated evolutionary trajectories, we investigate three issues: i) whether the distribution of connectivities changes between PPI subnetworks of essential and nonessential genes; ii) how gene conservation, measured both by the evolutionary retention index (ERI) and by evolutionary pressures, is related to the the connectivity of the corresponding protein; iii) how PPI connectivities are modulated by evolutionary and functional relationships, as represented by the Clusters of Orthologous Genes (COGs). We show that conservation, essentiality and functional specialisation of genes constrain the connectivity of the corresponding proteins in bacterial PPI networks. In particular, we isolate a core of highly connected proteins (with connectivities $k\ge40$), which is ubiquitous among the species considered here -- though mostly visible in the degree distributions of bacteria with small genomes (less than 1000 genes). The genes that belong to this highly connected core are conserved, essential and, in most cases, belong to the COG cluster J, related to ribosomal functions and to the processing of genetic information.

q-bio.GN↗

Essentiality, conservation, evolutionary pressure and codon bias in bacterial genes

Essential genes constitute the core of genes which cannot be mutated too much nor lost along the evolutionary history of a species. Natural selection is expected to be stricter on essential genes and on conserved (highly shared) genes, than on genes that are either nonessential or peculiar to a single or a few species. In order to further assess this expectation, we study here how essentiality of a gene is connected with its degree of conservation among several unrelated bacterial species, each one characterised by its own codon usage bias. Confirming previous results on E. coli, we show the existence of a universal exponential relation between gene essentiality and conservation in bacteria. Moreover we show that, within each bacterial genome, there are at least two groups of functionally distinct genes, characterised by different levels of conservation and codon bias: i) a core of essential genes, mainly related to cellular information processing; ii) a set of less conserved nonessential genes with prevalent functions related to metabolism. In particular, the genes in the first group are more retained among species, are subject to a stronger purifying conservative selection and display a more limited repertoire of synonymous codons. The core of essential genes is close to the minimal bacterial genome, which is in the focus of recent studies in synthetic biology, though we confirm that orthologs of genes that are essential in one species are not necessarily essential in other species. We also list a set of highly shared genes which, reasonably, could constitute a reservoir of targets for new anti-microbial drugs.

q-bio.GN↗

A fluctuation theorem for time-series of signal-response models with the backward transfer entropy

The irreversibility of trajectories in stochastic dynamical systems is linked to the structure of their causal representation in terms of Bayesian networks. We consider stochastic maps resulting from a time discretization with interval τof signal-response models, and we find an integral fluctuation theorem that sets the backward transfer entropy as a lower bound to the conditional entropy production. We apply this to a linear signal-response model providing analytical solutions, and to a nonlinear model of receptor-ligand systems. We show that the observational time τhas to be fine-tuned for an efficient detection of the irreversibility in time-series.

physics.data-an↗

Variants of intrinsic disorder in the human proteome

In this paper we propose a straightforward operational definition of variants of disordered proteins, taking the human proteome as a case study. The focus is on a distinction between mostly unstructured proteins and proteins which contain long unstructured regions accommodated in an overall folded structure. In particular we distinguish: i) Not disordered proteins (NDPs), that either have all their residues ordered or do not have disordered segments longer than 30 residues nor more than 30% of disordered residues; ii) Proteins with intrinsically disordered regions (IDRPs), that have at least one disordered domain longer than 30 residues, but disordered in less than 30% of their residues; iii) Proteins that are intrinsically disordered (IDPs), that have both at least one disordered segment longer than 30 residues and that are disordered on more than 30% of their residues; iv) Proteins with fragmented disorder (FRAG_IDPs), that do not have a disordered fragment longer than 30 residues but that, nevertheless, have at least 30% or more of their residues predicted as disordered. The potential use of these variants is checked over several groups of disease-related proteins. Our main conclusions point out that IDRPs are more similar to NDPs than to IDPs. IDRPs and NDPs have a similar functional repertoire and probably share a lock-and-key mechanism of interaction with substrates. IDRPs and IDPs are differently present among human disease-related proteins. IDRPs probably do not play a specific role in the development of complex diseases, since their frequency is similar in disease-related proteins and in the entire human proteome. IDPs can play a role in the emergence of cancer, neurodegenerative, thyroid and liver disease.

q-bio.BM↗

Causal influence in linear response models

The intuition of causation is so fundamental that almost every research study in life sciences refers to this concept. However a widely accepted formal definition of causal influence between observables is still missing. In the framework of linear Langevin networks without feedbacks (linear response models) we developed a measure of causal influence based on a decomposition of information flows over time. We discuss its main properties and compare it with other information measures like the Transfer Entropy. Finally we outline some difficulties of the extension to a general definition of causal influence for complex systems.

stat.OT↗

Codon Bias Patterns of $E.coli$'s Interacting Proteins

Synonymous codons, i.e., DNA nucleotide triplets coding for the same amino acid, are used differently across the variety of living organisms. The biological meaning of this phenomenon, known as codon usage bias, is still controversial. In order to shed light on this point, we propose a new codon bias index, $CompAI$, that is based on the competition between cognate and near-cognate tRNAs during translation, without being tuned to the usage bias of highly expressed genes. We perform a genome-wide evaluation of codon bias for $E.coli$, comparing $CompAI$ with other widely used indices: $tAI$, $CAI$, and $Nc$. We show that $CompAI$ and $tAI$ capture similar information by being positively correlated with gene conservation, measured by ERI, and essentiality, whereas, $CAI$ and $Nc$ appear to be less sensitive to evolutionary-functional parameters. Notably, the rate of variation of $tAI$ and $CompAI$ with ERI allows to obtain sets of genes that consistently belong to specific clusters of orthologous genes (COGs). We also investigate the correlation of codon bias at the genomic level with the network features of protein-protein interactions in $E.coli$. We find that the most densely connected communities of the network share a similar level of codon bias (as measured by $CompAI$ and $tAI$). Conversely, a small difference in codon bias between two genes is, statistically, a prerequisite for the corresponding proteins to interact. Importantly, among all codon bias indices, $CompAI$ turns out to have the most coherent distribution over the communities of the interactome, pointing to the significance of competition among cognate and near-cognate tRNAs for explaining codon usage adaptation.

q-bio.GN↗

On the abundance of intrinsically disordered proteins in the human proteome and its relation to diseases: there is no enrichment

Intrinsically disordered proteins are fascinating the community of protein science since the last decade, at least. There is a well-established line of research that intends to reveal the crucial role played by intrinsically disordered proteins (IDPs) in the development of human diseases. The main argument is that IDPs are differentially more present in groups of disease-related proteins. In this note we compare the frequency of disorder in human proteins, both disease-related and not. The frequency of disorder is comparable in the two sub-groups of proteins. Disorder is widespread in human proteins, but it is not a specific pre-requisite of proteins involved in the development of cancer, cardiovascular diseases, diabetes and neurodegenerative diseases. A tendency of cancer-related proteins to be statistically more disordered than the rest of human proteins is confirmed.

q-bio.BM↗

Amino acid composition and thermal stability of protein structures: the free energy geography of the Protein Data Bank

We study the combined influence of amino acid composition and chain length on the thermal stability of protein structures. A new parameterization of the internal free energy is considered, as the sum of hydrophobic effect, hydrogen-bond and de-hydration energy terms. We divided a non-redundant selection of protein structures from the Protein Data Bank into three groups: i) rich in order-promoting residues (OPR proteins); ii) rich in disorder-promoting residues (DPR proteins); iii) belonging to a twilight zone (TZ proteins). We observe a partition of PDB in several groups with different internal free energies, amino acid compositions and protein lengths. Internal free energy of 96% of the proteins analyzed ranges from -2 to -6.5 kJ/mol/res. We found many DPR and OPR proteins with the same relative thermal stability. Only OPR proteins with internal energy between -4 and -6.5 kJ/mol/res are observed to have chains longer than 200 residues, with a high de-hydration energy compensated by the hydrophobic effect. DPR and TZ proteins are shorter than 200 residues and they have an internal energy above -4 kJ/mol/res, with a few exceptions among TZ proteins. Hydrogen-bonds play an important role in the stabilization of these DPR folds, often higher than contact energy. The new parameterization of internal free energy let emerge a geography of thermal stabilities of PDB structures. Amino acid composition per se is not sufficient to determine the stability of protein folds, since. DPR and TZ proteins generally have a relatively high internal free energy, and they are stabilized by hydrogen-bonds. Long DPR proteins are not observed in the PDB, because their low hydrophobicity cannot compensate the high de-hydration energy necessary to accommodate residues within a highly packed globular fold.

q-bio.GN↗

On the performance of combined dichotomic predictors of natively unfolded proteins

The performance of single folding predictors and combination scores is critically evaluated. We test mean packing, mean pairwise energy and the new index gVSL2 on a dataset of 743 folded proteins and 81 natively unfolded proteins. These predictors have an individual performance comparable or even better than other proposed methods. We introduce here a strictly unanimous score S_{SU} that combines them but leaves undecided those sequences differently classified by two single predictors. The performance of the single predictors on a dataset purged from the proteins left unclassified by S_{SU}, significantly increases, indicating that unclassified proteins are mainly false predictions. Amino acid composition is the main determinant considered by these predictors, therefore unclassified proteins have a composition compatible with both folded and unfolded status. This is why purging a dataset from these ambiguous proteins increases the performance of single predictors. The percentage of proteins predicted as natively unfolded by S_{SU} in the three kingdoms are: 4.1% for Bacteria, 1.0% for Archaea and 20.0% for Eukarya; compatible with previous determinations. Evidence is given of a scaling law relating the number of natively unfolded proteins with the total number of proteins in a genome; a first estimate of the critical exponent is 1.95 +- 0.21

q-bio.BM↗

Is the unfoldome widespread in proteomes?

The term unfoldome has been recently used to indicate the universe of intrinsically disordered proteins. These proteins are characterized by an ensemble of high-flexible interchangeable conformations and therefore they can interact with many targets without requiring pre-existing stereo-chemical complementarity. It has been suggested that intrinsically disordered proteins are frequent in proteomes and disorder is widespread also in structured proteins. However, several studies raise some doubt about these views. It this paper we estimate the frequency of intrinsically disordered proteins in several living organisms by using the ratio S between the likelihood, for a protein sequence, of being composed mainly by order-promoting or disorder-promoting residues. We scan several proteomes from Archaea, Bacteria and Eukarya. We find the following figures: 1.63% for Archaea, 3.91% for Bacteria, 16.35% for Eukarya. The frequencies we found can be considered an upper bound to the real frequency of intrinsically disordered proteins in proteomes. Our estimates are lower than those previously reported in several studies. A scanning of proteins in the Protein Data Bank (PDB) searching for segments of non-observed residues reveals that segments of non-observed residues longer than 30 amino acids, are rare. Our observations support the idea that the spread of the unfoldome has been often overestimated. If we exclude some exceptions, the structure-function paradigm is generally valid and pre-existing stereo-chemical complementarity among structures remains an important requisite for interactions between biological macromolecules.

q-bio.GN↗

Combining predictors of natively unfolded proteins to detect a twilight zone between order and disorder in generic datasets

Natively unfolded proteins lack a well defined three dimensional structure but have important biological functions, suggesting a re-assignment of the structure-function paradigm. Many proteins have amino acidic compositions compatible both with the folded and unfolded status, and belong to a twilight zone between order and disorder. This makes difficult a dichotomic classification of protein sequences into folded and natively unfolded ones. In this methodological paper dichotomic folding indexes are considered: hydrophobicity-charge, mean packing, mean pairwise energy, Poodle-W and a new global index, that is called here gVSL2, based on the local disorder predictor VSL2. The performance of these indexes is evaluated on different datasets. Poodle-W, gVSL2 and mean pairwise energy have good performance and stability in all the datasets considered and are combined into a strictly unanimous combination score SSU, that leaves proteins unclassified when the consensus of all combined indexes is not reached. The unclassified proteins: i) belong to an overlap region in the vector space of amino acidic compositions occupied by both folded and unfolded proteins; ii) are composed by approximately the same number of order-promoting and disorder-promoting amino acids; iii) have a mean flexibility intermediate between that of folded and that of unfolded proteins. These proteins reasonably have physical properties intermediate between those of folded and those of natively unfolded proteins and their structural properties and evolutionary history are worth to be investigated.

q-bio.BM↗

Dynamical phase transitions in long-range Hamiltonian systems and Tsallis distributions with a time-dependent index

We study dynamical phase transitions in systems with long-range interactions, using the Hamiltonian Mean Field (HMF) model as a simple example. These systems generically undergo a violent relaxation to a quasi-stationary state (QSS) before relaxing towards Boltzmann equilibrium. In the collisional regime, the out-of-equilibrium one-particle distribution function (DF) is a quasi-stationary solution of the Vlasov equation, slowly evolving in time due to finite $N$ effects. For subcritical energies $7/12<U<3/4$, we exhibit cases where the DF is well-fitted by a Tsallis $q$-distribution with an index $q(t)$ slowly decreasing in time from $q\simeq 3$ (semi-ellipse) to $q=1$ (Boltzmann). When the index $q(t)$ reaches a critical value $q_{crit}(U)$, the non-magnetized (homogeneous) phase becomes Vlasov unstable and a dynamical phase transition is triggered, leading to a magnetized (inhomogeneous) state. While Tsallis distributions play an important role in our study, we explain this dynamical phase transition by using only conventional statistical mechanics. For supercritical energies, we report for the first time the existence of a magnetized QSS with a very long lifetime.

cond-mat.stat-mech↗

Number of natively unfolded proteins scales with genome size

Natively unfolded proteins exist as an ensemble of flexible conformations lacking a well defined tertiary structure along a large portion of their polypeptide chain. Despite the absence of a stable configuration, they are involved in important cellular processes. In this work we used from three indicators of folding status, derived from the analysis of mean packing and mean contact energy of a protein sequence as well as from VSL2, a disorder predictor, and we combined them into a consensus score to identify natively unfolded proteins in several genomes from Archaea, Bacteria and Eukarya. We found a high correlation among the number of predicted natively unfolded proteins and the number of proteins in the genomes. More specifically, the number of natively unfolded proteins scaled with the number of proteins in the genomes, with exponent 1.81 +- 0.10. This scaling law may be important to understand the relation between the number of natively unfolded proteins and their roles in cellular processes.

q-bio.GN↗

Long time behavior of quasi-stationary states of the Hamiltonian Mean-Field model

The Hamiltonian Mean-Field model has been investigated, since its introduction about a decade ago, to study the equilibrium and dynamical properties of long-range interacting systems. Here we study the long-time behavior of long-lived, out-of-equilibrium, quasi-stationary dynamical states, whose lifetime diverges in the thermodynamic limit. The nature of these states has been the object of a lively debate, in the recent past. We introduce a new numerical tool, based on the fluctuations of the phase of the instantaneous magnetization of the system. Using this tool, we study the quasi-stationary states that arise when the system is started from different classes of initial conditions, showing that the new observable can be exploited to compute the lifetime of these states. We also show that quasi-stationary states are present not only below, but also above the critical temperature of the second order magnetic phase transition of the model. We find that at supercritical temperatures the lifetime is much larger than at subcritical temperatures.

cond-mat.stat-mech↗