SearcharxivSearch

arXiv subjects

Terence Hwa

Publications and source records attributed to Terence Hwa.

At least 19 recordsLinked to original sources

Dynamic coexistence driven by physiological transitions in microbial communities

Microbial ecosystems are commonly modeled by fixed interactions between species in steady exponential growth states. However, microbes often modify their environments so strongly that they are forced out of the exponential state into stressed or non-growing states. Such dynamics are typical of ecological succession in nature and serial-dilution cycles in the laboratory. Here, we introduce a phenomenological model, the Community State model, to gain insight into the dynamic coexistence of microbes due to changes in their physiological states. Our model bypasses specific interactions (e.g., nutrient starvation, stress, aggregation) that lead to different combinations of physiological states, referred to collectively as "community states", and modeled by specifying the growth preference of each species along a global ecological coordinate, taken here to be the total community biomass density. We identify three key features of such dynamical communities that contrast starkly with steady-state communities: increased tolerance of community diversity to fast growth rates of species dominating different community states, enhanced community stability through staggered dominance of different species in different community states, and increased requirement on growth dominance for the inclusion of late-growing species. These features, derived explicitly for simplified models, are proposed here to be principles aiding the understanding of complex dynamical communities. Our model shifts the focus of ecosystem dynamics from bottom-up studies based on idealized inter-species interaction to top-down studies based on accessible macroscopic observables such as growth rates and total biomass density, enabling quantitative examination of community-wide characteristics.

q-bio.PE

Constrained Allocation Flux Balance Analysis

New experimental results on bacterial growth inspire a novel top-down approach to study cell metabolism, combining mass balance and proteomic constraints to extend and complement Flux Balance Analysis. We introduce here Constrained Allocation Flux Balance Analysis, CAFBA, in which the biosynthetic costs associated to growth are accounted for in an effective way through a single additional genome-wide constraint. Its roots lie in the experimentally observed pattern of proteome allocation for metabolic functions, allowing to bridge regulation and metabolism in a transparent way under the principle of growth-rate maximization. We provide a simple method to solve CAFBA efficiently and propose an "ensemble averaging" procedure to account for unknown protein costs. Applying this approach to modeling E. coli metabolism, we find that, as the growth rate increases, CAFBA solutions cross over from respiratory, growth-yield maximizing states (preferred at slow growth) to fermentative states with carbon overflow (preferred at fast growth). In addition, CAFBA allows for quantitatively accurate predictions on the rate of acetate excretion and growth yield based on only 3 parameters determined by empirical growth laws.

q-bio.MN

Stripe formation in bacterial systems with density-suppressed motility

Engineered bacteria in which motility is reduced by local cell density generate periodic stripes of high and low density when spotted on agar plates. We study theoretically the origin and mechanism of this process in a kinetic model that includes growth and density-suppressed motility of the cells. The spreading of a region of immotile cells into an initially cell-free region is analyzed. From the calculated front profile we provide an analytic ansatz to determine the phase boundary between the stripe and the no-stripe phases. The influence of various parameters on the phase boundary is discussed.

cond-mat.soft

Direct-coupling analysis of residue co-evolution captures native contacts across many protein families

The similarity in the three-dimensional structures of homologous proteins imposes strong constraints on their sequence variability. It has long been suggested that the resulting correlations among amino acid compositions at different sequence positions can be exploited to infer spatial contacts within the tertiary protein structure. Crucial to this inference is the ability to disentangle direct and indirect correlations, as accomplished by the recently introduced Direct Coupling Analysis (DCA) (Weigt et al. (2009) Proc Natl Acad Sci 106:67). Here we develop a computationally efficient implementation of DCA, which allows us to evaluate the accuracy of contact prediction by DCA for a large number of protein domains, based purely on sequence information. DCA is shown to yield a large number of correctly predicted contacts, recapitulating the global structure of the contact map for the majority of the protein domains examined. Furthermore, our analysis captures clear signals beyond intra- domain residue contacts, arising, e.g., from alternative protein conformations, ligand- mediated residue couplings, and inter-domain interactions in protein oligomers. Our findings suggest that contacts predicted by DCA can be used as a reliable guide to facilitate computational predictions of alternative protein conformations, protein complex formation, and even the de novo prediction of protein domain structures, provided the existence of a large number of homologous sequences which are being rapidly made available due to advances in genome sequencing.

q-bio.QM

Dissecting the Specificity of Protein-Protein Interaction in Bacterial Two-Component Signaling: Orphans and Crosstalks

Predictive understanding of the myriads of signal transduction pathways in a cell is an outstanding challenge of systems biology. Such pathways are primarily mediated by specific but transient protein-protein interactions, which are difficult to study experimentally. In this study, we dissect the specificity of protein-protein interactions governing two-component signaling (TCS) systems ubiquitously used in bacteria. Exploiting the large number of sequenced bacterial genomes and an operon structure which packages many pairs of interacting TCS proteins together, we developed a computational approach to extract a molecular interaction code capturing the preferences of a small but critical number of directly interacting residue pairs. This code is found to reflect physical interaction mechanisms, with the strongest signal coming from charged amino acids. It is used to predict the specificity of TCS interaction: Our results compare favorably to most available experimental results, including the prediction of 7 (out of 8 known) interaction partners of orphan signaling proteins in Caulobacter crescentus. Surveying among the available bacterial genomes, our results suggest 15~25% of the TCS proteins could participate in out-of-operon "crosstalks". Additionally, we predict clusters of crosstalking candidates, expanding from the anecdotally known examples in model organisms. The tools and results presented here can be used to guide experimental studies towards a system-level understanding of two-component signaling.

q-bio.MN

Designing sequential transcription logic: a simple genetic circuit for conditional memory

The ability to learn and respond to recurrent events depends on the capacity to remember transient biological signals received in the past. Moreover, it may be desirable to remember or ignore these transient signals conditioned upon other signals that are active at specific points in time or in unique environments. Here, we propose a simple genetic circuit in bacteria that is capable of conditionally memorizing a signal in the form of a transcription factor concentration. The circuit behaves similarly to a "data latch" in an electronic circuit, i.e. it reads and stores an input signal only when conditioned to do so by a "read command". Our circuit is of the same size as the well-known genetic toggle switch (an unconditional latch) which consists of two mutually repressing genes, but is complemented with a "regulatory front end" involving protein heterodimerization as a simple way to implement conditional control. Deterministic and stochastic analysis of the circuit dynamics indicate that an experimental implementation is feasible based on well-characterized genes and proteins. It is not known, to which extent molecular networks are able to conditionally store information in natural contexts for bacteria. However, our results suggest that such sequential logic elements may be readily implemented by cells through the combination of existing protein-protein interactions and simple transcriptional regulation.

q-bio.MN

Growth-rate dependent partitioning of RNA polymerases in bacteria

Physiological changes which result in changes in bacterial gene expression are often accompanied by changes in the growth rate for fast adapting enteric bacteria. Since the availability of RNA polymerase (RNAP) in cells is dependent on the growth rate, transcriptional control involves not only the regulation of promoters, but also depends on the available (or free) RNAP concentration which is difficult to quantify directly. Here we develop a simple physical model describing the partitioning of cellular RNAP into different classes: RNAPs transcribing mRNA and ribosomal RNA (rRNA), RNAPs non-specifically bound to DNA, free RNAP, and immature RNAP. Available experimental data for E. coli allow us to determine the two unknown parameters of the model and hence deduce the free RNAP concentration at different growth rates. The results allow us to predict the growth-rate dependence of the activities of constitutive (unregulated) promoters, and to disentangle the growth-rate dependent regulation of promoters (e.g., the promoters of rRNA operons) from changes in transcription due to changes in the free RNAP concentration at different growth rates. Our model can quantitatively account for the observed changes in gene expression patterns in mutant E. coli strains with altered levels of RNAP expression without invoking additional parameters. Applying our model to the case of the stringent response following amino acid starvation, we can evaluate the plausibility of various scenarios of passive transcriptional control proposed to account for the observed changes in the expression of rRNA and biosynthetic operons.

q-bio.SC

Stochasticity and traffic jams in the transcription of ribosomal RNA: Intriguing role of termination and antitermination

In fast growing bacteria, ribosomal RNA (rRNA) is required to be transcribed at very high rates to sustain the high cellular demand on ribosome synthesis. This results in dense traffic of RNA polymerases (RNAP). We developed a stochastic model, integrating results of single-molecule and quantitative in vivo studies of E. coli, to evaluate the quantitative effect of pausing, termination, and antitermination on rRNA transcription. Our calculations reveal that in dense RNAP traffic, spontaneous pausing of RNAP can lead to severe "traffic jams", as manifested in the broad distribution of inter-RNAP distances and can be a major factor limiting transcription and hence growth. Our results suggest the suppression of these pauses by the ribosomal antitermination complex to be essential at fast growth. Moreover, unsuppressed pausing by even a few non-antiterminated RNAPs can already reduce transcription drastically under dense traffic. However, the termination factor Rho can remove the non-antiterminated RNAPs and restore fast transcription. The results thus suggest an intriguing role by Rho to enhance rather than attenuate rRNA transcription.

q-bio.SC

Deterministic characterization of stochastic genetic circuits

For cellular biochemical reaction systems where the numbers of molecules is small, significant noise is associated with chemical reaction events. This molecular noise can give rise to behavior that is very different from the predictions of deterministic rate equation models. Unfortunately, there are few analytic methods for examining the qualitative behavior of stochastic systems. Here we describe such a method that extends deterministic analysis to include leading-order corrections due to the molecular noise. The method allows the steady-state behavior of the stochastic model to be easily computed, facilitates the mapping of stability phase diagrams that include stochastic effects and reveals how model parameters affect noise susceptibility, in a manner not accessible to numerical simulation. By way of illustration we consider two genetic circuits: a bistable positive-feedback loop and a negative-feedback oscillator. We find in the positive feedback circuit that translational activation leads to a far more stable system than transcriptional control. Conversely, in a negative-feedback loop triggered by a positive-feedback switch, the stochasticity of transcriptional control is harnessed to generate reproducible oscillations.

q-bio.MN

Stochastic fluctuations in metabolic pathways

Fluctuations in the abundance of molecules in the living cell may affect its growth and well being. For regulatory molecules (e.g., signaling proteins or transcription factors), fluctuations in their expression can affect the levels of downstream targets in a network. Here, we develop an analytic framework to investigate the phenomenon of noise correlation in molecular networks. Specifically, we focus on the metabolic network, which is highly inter-linked, and noise properties may constrain its structure and function. Motivated by the analogy between the dynamics of a linear metabolic pathway and that of the exactly soluable linear queueing network or, alternatively, a mass transfer system, we derive a plethora of results concerning fluctuations in the abundance of intermediate metabolites in various common motifs of the metabolic network. For all but one case examined, we find the steady-state fluctuation in different nodes of the pathways to be effectively uncorrelated. Consequently, fluctuations in enzyme levels only affect local properties and do not propagate elsewhere into metabolic networks, and intermediate metabolites can be freely shared by different reactions. Our approach may be applicable to study metabolic networks with more complex topologies, or protein signaling networks which are governed by similar biochemical reactions. Possible implications for bioinformatic analysis of metabolimic data are discussed.

q-bio.MN

Identification and Measurement of Neighbor Dependent Nucleotide Substitution Processes

The presence of neighbor dependencies generated a specific pattern of dinucleotide frequencies in all organisms. Especially, the CpG-methylation-deamination process is the predominant substitution process in vertebrates and needs to be incorporated into a more realistic model for nucleotide substitutions. Based on a general framework of nucleotide substitutions we develop a method that is able to identify the most relevant neighbor dependent substitution processes, measure their strength, and judge their importance to be included into the modeling. Starting from a model for neighbor independent nucleotide substitution we successively add neighbor dependent substitution processes in the order of their ability to increase the likelihood of the model describing given data. The analysis of neighbor dependent nucleotide substitutions in human, zebrafish and fruit fly is presented. A web server to perform the presented analysis is publicly available.

q-bio.GN

Substantial regional variation in substitution rates in the human genome: importance of GC content, gene density and telomere-specific effects

This study presents the first global, 1 Mbp level analysis of patterns of nucleotide substitutions along the human lineage. The study is based on the analysis of a large amount of repetitive elements deposited into the human genome since the mammalian radiation, yielding a number of results that would have been difficult to obtain using the more conventional comparative method of analysis. This analysis revealed substantial and consistent variability of rates of substitution, with the variability ranging up to 2-fold among different regions. The rates of substitutions of C or G nucleotides with A or T nucleotides vary much more sharply than the reverse rates suggesting that much of that variation is due to differences in mutation rates rather than in the probabilities of fixation of C/G vs. A/T nucleotides across the genome. For all types of substitution we observe substantially more hotspots than coldspots, with hotspots showing substantial clustering over tens of Mbp's. Our analysis revealed that GC-content of surrounding sequences is the best predictor of the rates of substitution. The pattern of substitution appears very different near telomeres compared to the rest of the genome and cannot be explained by the genome-wide correlations of the substitution rates with GC content or exon density. The telomere pattern of substitution is consistent with natural selection or biased gene conversion acting to increase the GC-content of the sequences that are within 10-15 Mbp away from the telomere.

q-bio.GN

Nonlinear Protein Degradation and the Function of Genetic Circuits

The functions of most genetic circuits require sufficient degrees of cooperativity in the circuit components. While mechanisms of cooperativity have been studied most extensively in the context of transcriptional initiation control, cooperativity from other processes involved in the operation of the circuits can also play important roles. In this study, we examine a simple kinetic source of cooperativity stemming from the nonlinear degradation of multimeric proteins. Ample experimental evidence suggests that protein subunits can degrade less rapidly when associated in multimeric complexes, an effect we refer to as cooperative stability. For dimeric transcription factors, this effect leads to a concentration-dependence in the degradation rate because monomers, which are predominant at low concentrations, will be more rapidly degraded. Thus cooperative stability can effectively widen the accessible range of protein levels in vivo. Through theoretical analysis of two exemplary genetic circuits in bacteria, we show that such an increased range is important for the robust operation of genetic circuits as well as their evolvability. Our calculations demonstrate that a few-fold difference between the degradation rate of monomers and dimers can already enhance the function of these circuits substantially. These results suggest that cooperative stability needs to be considered explicitly and characterized quantitatively in any systematic experimental or theoretical study of gene circuits.

q-bio.MN

Transcriptional Regulation by the Numbers 1: Models

The study of gene regulation and expression is often discussed in quantitative terms. In particular, the expression of genes is regularly characterized with respect to how much, how fast, when and where. Whether discussing the level of gene expression in a bacterium or its precise location within a developing embryo, the natural language for these experiments is that of numbers. Such quantitative data demands quantitative models. We review a class of models ("thermodynamic models") which exploit statistical mechanics to compute the probability that RNA polymerase is at the appropriate promoter. This provides a mathematically precise elaboration of the idea that activators are agents of recruitment which increase the probability that RNA polymerase will be found at the promoter of interest. We discuss a framework which describes the interactions of repressors, activators, helper molecules and RNA polymerase using the concept of effective concentrations, expressed in terms of a function we call the "regulation factor". This analysis culminates in an expression for the probability of RNA polymerase binding at the promoter of interest as a function of the number of regulatory proteins in the cell. In a companion paper [1], these ideas are applied to several case studies which illustrate the use of the general formalism.

q-bio.MN

Transcriptional Regulation by the Numbers 2: Applications

With the increasing amount of experimental data on gene expression and regulation, there is a growing need for quantitative models to describe the data and relate them to the different contexts. The thermodynamic models reviewed in the preceding paper provide a useful framework for the quantitative analysis of bacterial transcription regulation. We review a number of well-characterized bacterial promoters that are regulated by one or two species of transcription factors, and apply the thermodynamic framework to these promoters. We show that the framework allows one to quantify vastly different forms of gene expression using a few parameters. As such, it provides a compact description useful for higher-level studies, e.g., of genetic networks, without the need to invoke the biochemical details of every component. Moreover, it can be used to generate hypotheses on the likely mechanisms of transcriptional control.

q-bio.MN

Translocation of structured polynucleotides through nanopores

We investigate theoretically the translocation of structured RNA/DNA molecules through narrow pores which allow single but not double strands to pass. The unzipping of basepaired regions within the molecules presents significant kinetic barriers for the translocation process. We show that this circumstance may be exploited to determine the full basepairing pattern of polynucleotides, including RNA pseudoknots. The crucial requirement is that the translocation dynamics (i.e., the length of the translocated molecular segment) needs to be recorded as a function of time with a spatial resolution of a few nucleotides. This could be achieved, for instance, by applying a mechanical driving force for translocation and recording force-extension curves (FEC's) with a device such as an atomic force microscope or optical tweezers. Our analysis suggests that with this added spatial resolution, nanopores could be transformed into a powerful experimental tool to study the folding of nucleic acids.

cond-mat.soft

DNA Sequence Evolution with Neighbor-Dependent Mutation

We introduce a model of DNA sequence evolution which can account for biases in mutation rates that depend on the identity of the neighboring bases. An analytic solution for this class of non-equilibrium models is developed by adopting well-known methods of nonlinear dynamics. Results are presented for the CpG-methylation-deamination process which dominates point substitutions in vertebrates. The dinucleotide frequencies generated by the model (using empirically obtained mutation rates) match the overall pattern observed in non-coding DNA. A web-based tool has been constructed to compute single- and dinucleotide frequencies for arbitrary neighbor-dependent mutation rates. Alsoprovided is the backward procedure to infer the mutation rates using maximum likelihood analysis given the observed single- and dinucleotide frequencies. Reasonable estimates of the mutation rates can be obtained very efficiently, using generic non-coding DNA sequences as input, after masking outlong homonucleotide subsequences. Our method is much more convenient and versatile to use than the traditional method of deducing mutation rates by counting mutation events in carefully chosen sequences. More generally, our approach provides a more realistic but still tractable description of non-coding genomic DNA, and may be used as a null model for various sequence analysis applications.

physics.bio-ph

Distinct changes of genomic biases in nucleotide substitution at the time of mammalian radiation

Differences in the regional substitution patterns in the human genome created patterns of large-scale variation of base composition known as genomic isochores. To gain insight into the origin of the genomic isochores we develop a maximum likelihood approach to determine the history of substitution patterns in the human genome. This approach utilizes the vast amount of repetitive sequence deposited in the human genome over the past ~250 MYR. Using this approach we estimate the frequencies of seven types of substitutions: the four transversions, two transitions, and the methyl-assisted transition of cytosine in CpG. Comparing substitutional patterns in repetitive elements of various ages, we reconstruct the history of the base-substitutional process in the different isochores for the past 250 Myr. At around 90 Myr ago (around the time of the mammalian radiation), we find an abrupt 4- to 8-fold increase of the cytosine transition rate in CpG pairs compared to that of the reptilian ancestor. Further analysis of nucleotide substitutions in regions with different GC-content reveals concurrent changes in the substitutional patterns. While the substitutional pattern was dependent on the regional GC-content in such ways that it preserved the regional GC-content before the mammalian radiation, it lost this dependence afterwards. The substitutional pattern changed from an isochore-preserving to an isochore-degrading one. We conclude that isochores have been established before the radiation of the eutherian mammals and have been subject to the process of homogenization since then.

cond-mat.stat-mech