SearcharxivSearch

arXiv subjects

Zhanshan

Publications and source records attributed to Zhanshan.

2 recordsLinked to original sources

Individual-Level SNP Diversity and Similarity Profiles

Classic concepts of genetic (gene) diversity (heterozygosity) such as Nei (1973: PNAS) and Nei and Li (1979: PNAS) nucleotide diversity were defined within the context of populations. Although variations are often measured in population context, the basic carriers of variation are individuals. Hence, measuring variations such as SNP of individual against a reference genome, which has been ignored currently, is certainly of its own right. Indeed, similar practice has been a tradition in ecology, where the basic framework of diversity measure is individual community sample. We propose to use Renyi-entropy-derived Hill numbers to define SNP (single nucleotide polymorphism) diversity (including alpha-, beta-, and gamma-diversities) and similarity profiles. Hill numbers are derived from Renyi entropy, of which Shannon entropy is a special case and which have found widely applications including measuring the quantum information entanglement, wealth distribution in economics and ecological diversity. The newly proposed SNP diversity not only complements the existing genetic diversity concepts by offering individual-level metrics, but also offers building blocks for comparative genetic analysis at higher levels. The profile concept also helps to resolve a dilemma in measuring diversity: the choice from various diversity indexes, because diversity profile unifies some of the most commonly used indexes (as special cases) with different diversity orders (along the rareness-commonness spectrum of gene mutations). Finally, the profiles can be estimated with rarefaction approach, which may help to relieve some effect of insufficient sequencing coverage.

q-bio.PE

DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies

(An updated version of this manuscript has been accepted to Scientific Reports in 2016, please refer to http://www.nature.com/articles/srep31900) The highly anticipated transition from next generation sequencing (NGS) to third generation sequencing (3GS) has been difficult primarily due to high error rates and excessive sequencing cost. The high error rates make the assembly of long erroneous reads of large genomes challenging because existing software solutions are often overwhelmed by error correction tasks. Here we report a hybrid assembly approach that simultaneously utilizes NGS and 3GS data to address both issues. We gain advantages from three general and basic design principles: (i) Compact representation of the long reads lead to efficient alignments. (ii) Base-level errors can be skipped; structural errors need to be detected and corrected. (iii) Structurally correct 3GS reads are assembled and polished. In our implementation, preassembled NGS contigs are used to derive the compact representation of the long reads, which established an algorithmic conversion from a de Bruijn graph to an overlap graph, the two major assembly paradigms. Moreover, since NGS and 3GS data can compensate each other, our hybrid assembly approach reduces both of their sequencing requirements. Experiments show that our software is able to assemble mammalian-sized genomes orders of magnitude more efficiently in time than existing methods, while saving about half of the sequencing cost.

q-bio.GN