SearcharxivSearch

arXiv subjects

William W. Chen

Publications and source records attributed to William W. Chen.

4 recordsLinked to original sources

All-atom ab initio folding of a diverse set of proteins

Natural proteins fold to a unique, thermodynamically dominant state. Modeling of the folding process and prediction of the native fold of proteins are two major unsolved problems in biophysics. Here, we show successful all-atom ab initio folding of a representative diverse set of proteins, using a minimalist transferable energy model that consists of two-body atom-atom interactions, hydrogen-bonding, and a local sequence energy term that models sequence-specific chain stiffness. Starting from a random coil, the native-like structure was observed during replica exchange Monte Carlo (REMC) simulation for most proteins regardless of their structural classes; the lowest energy structure was close to native- in the range of 2-6 A root-mean-square deviation (RMSD). Our results demonstrate that the successful all-atom folding of a protein chain to its native state is governed by only a few crucial energetic terms.

q-bio.BM

Information and Protein Interfaces

To confer high specificity and affinity in binding, contacts at interfaces between two interacting macromolecules are expected to exhibit pair preferences for types of atoms or residues. Here we quantify these preferences by measuring the mutual information of contacts for 895 protein-protein interfaces. The information content is significant and is highest at the atomic resolution. A simple phenomenological theory reveals a connection between information at interfaces and the free energy spectrum of association. The connection is presented in the form of a relation between mutual information and the energy gap of the native bound state to off-target bound states. Measurement of information content in designed lattice interfaces show the predicted scaling behavior to the energy gap. Our theory also suggests that mutual information in contacts emerges by a selection mechanism, and that strong selection, or high conservation, of residues should lead to correspondingly high mutual information. Amino acids which contribute more heavily to information content are then expected to be more conserved. We verify this by showing a statistically significant correlation between the conservation of each of the twenty amino acids and their individual contribution to the information content at protein-protein interfaces

q-bio.BM

Entropic stabilization of proteins and its proteomic consequences

We report here a new entropic mechanism of protein thermostability due to residual dynamics of rotamer isomerization in native state. All-atom simulations show that Lysines have much greater number of accessible rotamers than Arginines in folded states of proteins. This finding suggests that Lysines would preferentially entropically stabilize the native state. Indeed we show in computational experiments that Arginine-to-Lysine amino acid substitutions result in noticeable stabilization of proteins. We then hypothesize that if evolution uses this physical mechanisms in its strategies of thermophilic adaptation then hyperthermostable organisms would have much greater content of Lysines in their proteomes than of comparable in size and similarly charged Arginines.. Consistent with that, high-throughput comparative analysis of complete proteomes shows extremely strong bias towards Arginine-to-Lysine replacement in hyperthermophilic organisms and overall much greater content of Lysines than Arginines in hyperthermophiles. This finding cannot be explained by GC compositional biases. Our study provides an example of how analysis of a delicate physical mechanism of thermostability helps to resolve a puzzle in comparative genomics as to why aminoacid compositions of hyperthermophilic proteomes are significantly biased towards Lysines but not Arginines

q-bio.BM

An Exact Model of Fluctuations in Gene Expression

Fluctuations in the measured mRNA levels of unperturbed cells under fixed conditions have often been viewed as an impediment to the extraction of information from expression profiles. Here, we argue that such expression fluctuations should themselves be studied as a source of valuable information about the underlying dynamics of genetic networks. By analyzing microarray data taken from Saccharomyces cerevisiae, we demonstrate that correlations in expression fluctuations have a highly statistically significant dependence on gene function, and furthermore exhibit a remarkable scale-free network structure. We therefore present what we view to be the simplest phenomenological model of a genetic network which can account for the presence of biological information in transcript level fluctuations. We proceed to exactly solve this model using a path integral technique and derive several quantitative predictions. Finally, we propose several experiments by which these predictions might be rigorously tested.

q-bio.MN