Searcharxiv⌕ Search

arXiv subjects

Erik van Nimwegen

Publications and source records attributed to Erik van Nimwegen.

8 recordsLinked to original sources

Scaling laws in the functional content of genomes: Fundamental constants of evolution?

With the number of fully-sequenced genomes now well over a hundred it has become possible to start investigating if there are any quantitative regularities in the genetic make-up of genomes. In (physics/0307001), I originally showed that the numbers of genes in different functional categories scale as power laws in the total number of genes in the genome. In this chapter I revisit these results with more recent data and go into considerable more depth regarding the implications of these scaling laws for our understanding of the regulatory design of cells. In addition, I further develop the evolutionary model first proposed in (physics/0307001), which suggests that the exponents of the observed scaling laws correspond to fundamental constants of the evolutionary process. In particular, I put forward an hypothesis for the approximately quadratic scaling of regulatory and signal transducing genes with genome size.

q-bio.GN↗

Scaling laws in the functional content of genomes

With the number of sequenced genomes now over one hundred, and the availability of rough functional annotations for a substantial proportion of their genes, it has become possible to study the statistics of gene content across genomes. Here I show that, for many high-level functional categories, the number of genes in the category scales as a power-law in the total number of genes in the genome. The occurrence of such scaling laws can be explained with a simple theoretical model, and this model suggests that the exponents of the observed scaling laws correspond to universal constants of the evolutionary process. I discuss some consequences of these scaling laws for our understanding of organism design.

physics.bio-ph↗

Probabilistic Clustering of Sequences: Inferring new bacterial regulons by comparative genomics

Genome wide comparisons between enteric bacteria yield large sets of conserved putative regulatory sites on a gene by gene basis that need to be clustered into regulons. Using the assumption that regulatory sites can be represented as samples from weight matrices we derive a unique probability distribution for assignments of sites into clusters. Our algorithm, 'PROCSE' (probabilistic clustering of sequences), uses Monte-Carlo sampling of this distribution to partition and align thousands of short DNA sequences into clusters. The algorithm internally determines the number of clusters from the data, and assigns significance to the resulting clusters. We place theoretical limits on the ability of any algorithm to correctly cluster sequences drawn from weight matrices (WMs) when these WMs are unknown. Our analysis suggests that the set of all putative sites for a single genome (e.g. E. coli) is largely inadequate for clustering. When sites from different genomes are combined and all the homologous sites from the various species are used as a block, clustering becomes feasible. We predict 50-100 new regulons as well as many new members of existing regulons, potentially doubling the number of known regulatory sites in E. coli.

physics.bio-ph↗

Metastable Evolutionary Dynamics: Crossing Fitness Barriers or Escaping via Neutral Paths?

We analytically study the dynamics of evolving populations that exhibit metastability on the level of phenotype or fitness. In constant selective environments, such metastable behavior is caused by two qualitatively different mechanisms. One the one hand, populations may become pinned at a local fitness optimum, being separated from higher-fitness genotypes by a {\em fitness barrier} of low-fitness genotypes. On the other hand, the population may only be metastable on the level of phenotype or fitness while, at the same time, diffusing over {\em neutral networks} of selectively neutral genotypes. Metastability occurs in this case because the population is separated from higher-fitness genotypes by an {\em entropy barrier}: The population must explore large portions of these neutral networks before it discovers a rare connection to fitter phenotypes. We derive analytical expressions for the barrier crossing times in both the fitness barrier and entropy barrier regime. In contrast with ``landscape'' evolutionary models, we show that the waiting times to reach higher fitness depend strongly on the width of a fitness barrier and much less on its height. The analysis further shows that crossing entropy barriers is faster by orders of magnitude than fitness barrier crossing. Thus, when populations are trapped in a metastable phenotypic state, they are most likely to escape by crossing an entropy barrier, along a neutral path in genotype space. If no such escape route along a neutral path exists, a population is most likely to cross a fitness barrier where the barrier is {\em narrowest}, rather than where the barrier is shallowest.

adap-org↗

Neutral Evolution of Mutational Robustness

We introduce and analyze a general model of a population evolving over a network of selectively neutral genotypes. We show that the population's limit distribution on the neutral network is solely determined by the network topology and given by the principal eigenvector of the network's adjacency matrix. Moreover, the average number of neutral mutant neighbors per individual is given by the matrix spectral radius. This quantifies the extent to which populations evolve mutational robustness: the insensitivity of the phenotype to mutations. Since the average neutrality is independent of evolutionary parameters---such as, mutation rate, population size, and selective advantage---one can infer global statistics of neutral network topology using simple population data available from {\it in vitro} or {\it in vivo} evolution. Populations evolving on neutral networks of RNA secondary structures show excellent agreement with our theoretical predictions.

adap-org↗

The Evolutionary Unfolding of Complexity

We analyze the population dynamics of a broad class of fitness functions that exhibit epochal evolution---a dynamical behavior, commonly observed in both natural and artificial evolutionary processes, in which long periods of stasis in an evolving population are punctuated by sudden bursts of change. Our approach---statistical dynamics---combines methods from both statistical mechanics and dynamical systems theory in a way that offers an alternative to current ``landscape'' models of evolutionary optimization. We describe the population dynamics on the macroscopic level of fitness classes or phenotype subbasins, while averaging out the genotypic variation that is consistent with a macroscopic state. Metastability in epochal evolution occurs solely at the macroscopic level of the fitness distribution. While a balance between selection and mutation maintains a quasistationary distribution of fitness, individuals diffuse randomly through selectively neutral subbasins in genotype space. Sudden innovations occur when, through this diffusion, a genotypic portal is discovered that connects to a new subbasin of higher fitness genotypes. In this way, we identify innovations with the unfolding and stabilization of a new dimension in the macroscopic state space. The architectural view of subbasins and portals in genotype space clarifies how frozen accidents and the resulting phenotypic constraints guide the evolution to higher complexity.

adap-org↗

Optimizing Epochal Evolutionary Search: Population-Size Dependent Theory

Epochal dynamics, in which long periods of stasis in an evolving population are punctuated by a sudden burst of change, is a common behavior in both natural and artificial evolutionary processes. We analyze the population dynamics for a class of fitness functions that exhibit epochal behavior using a mathematical framework developed recently. In the latter the approximations employed led to a population-size independent theory that allowed us to determine optimal mutation rates. Here we extend this approach to include the destabilization of epochs due to finite-population fluctuations and show that this dynamical behavior often occurs around the optimal parameter settings for efficient search. The resulting, more accurate theory predicts the total number of fitness function evaluations to reach the global optimum as a function of mutation rate, population size, and the parameters specifying the fitness function. We further identify a generalized error threshold, smoothly bounding the two-dimensional regime of mutation rates and population sizes for which evolutionary search operates efficiently.

adap-org↗

Optimizing Epochal Evolutionary Search: Population-Size Independent Theory

Epochal dynamics, in which long periods of stasis in population fitness are punctuated by sudden innovations, is a common behavior in both natural and artificial evolutionary processes. We use a recent quantitative mathematical analysis of epochal evolution to estimate, as a function of population size and mutation rate, the average number of fitness function evaluations to reach the global optimum. This is then used to derive estimates of and bounds on evolutionary parameters that minimize search effort.

adap-org↗