SearcharxivSearch

arXiv subjects

Li-Ching Hsieh

Publications and source records attributed to Li-Ching Hsieh.

3 recordsLinked to original sources

Divergence and Shannon information in genomes

Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition. We find the following: SI (but not divergence) is inversely proportional to sequence length for a random sequence but is length-independent for genomes; the genomic SI is always greater and, for shorter words and longer sequences, hundreds to thousands times greater than the SI in a random sequence whose length and composition match those of the genome; genomic SIs appear to have word-length dependent universal values. The universality is inferred to be an evolution footprint of a universal mode for genome growth.

q-bio.GN

Quasireplicas and universal lengths of microbial genomes

Statistical analysis of distributions of occurrence frequencies of short words in 108 microbial complete genomes reveals the existence of a set of universal "root-sequence lengths" shared by all microbial genomes. These lengths and their universality give powerful clues to the way microbial genomes are grown. We show that the observed genomic properties are explained by a model for genome growth in which primitive genomes grew mainly by maximally stochastic duplications of short segments from an initial length of about 200 nucleotides (nt) to a length of about one million nt typical of microbial genomes. The relevance of the result of this study to the nature of simultaneous random growth and information acquisition by genomes, to the so-called RNA world in which life evolved before the rise of proteins and enzymes and to several other topics are discussed.

physics.bio-ph

Evidence for growth of microbial genomes by short segmental duplications

We show that textual analysis of microbial genomes reveal telling footprints of the early evolution of the genomes. The frequencies of word occurrence of random DNA sequences considered as texts in their four nucleotides are expected to obey Poisson distributions. It is noticed that for words less than nine letters the average width of the distributions for complete microbial genomes is many times that of a Poisson distribution. We interpret this phenomenon as follows: the genome is a large system that possesses the statistical characteristics of a much smaller ``random'' system, and certain textual statistical properties of genomes we now see are remnants of those of their ancestral genomes, which were much shorter than the genomes are now. This interpretation suggests a simple biologically plausible model for the growth of genomes: the genome first grows randomly to an initial length of approximately one thousand nucleotides (1k nt), or about one thousandth of its final length, thereafter mainly grows by random segmental duplication. We show that using duplicated segments averaging around 25 nt, the model sequences generated possess statistical properties characteristic of present day genomes. Both the initial length and the duplicated segment length support an RNA world at the time duplication began. Random segmental duplication would greatly enhance the ability of a genome to use its hard-to-acquire codes repeatedly, and a genome that practiced it would have evolved enormously faster than those that did not.

physics.bio-ph