SearcharxivSearch

arXiv subjects

Conrad J. Burden

Publications and source records attributed to Conrad J. Burden.

At least 19 recordsLinked to original sources

The Feller diffusion as the limit of a coalescent point process

The Feller diffusion is studied as the limit of a coalescent point process in which the density of the node height distribution is skewed towards zero. Using a unified approach, a number of recent results pertaining to scaling limits of branching processes are reviewed and reinterpreted as properties of the Feller diffusion arising from this limit. The notion of Bernoulli sampling of a finite population is extended to the diffusion limit to cover finite Poisson-distributed samples drawn from infinite continuum populations. We show that the coalescent tree of a Poisson-sampled Feller diffusion corresponds to a coalescent point process with a node height distribution taking the same algebraic form as that of a Bernoulli-sampled birth-death process. By adapting methods for analysing k-sampled birth-death processes, in which the sample size is pre-specified, we develop methods for studying the coalescent properties of the k-sampled Feller diffusion.

math.PR

The Feller diffusion conditioned on a single ancestral founder

We examine the distributional properties of a Feller diffusion $(X(τ))_{τ\in [0, t]}$ conditioned on the current population $X(t)$ having a single ancestor at time zero. The approach is novel and is based on an interpretation of Feller's original solution according to which the current population is comprised of a Poisson number of exponentially distributed families, each descended from a single ancestor. The distribution of the number of ancestors at intermediate times and the joint density of coalescent times is determined under assumptions of initiation of the process from a single ancestor at a specified time in the past, including infinitely far in the past, and for the case of a uniform prior on the time since initiation. Also calculated are the joint distribution of the time since the most recent common ancestor of the current population and the contemporaneous population size at that time under different assumptions on the time since initiation. In each case exact solutions are given for supercritical, critical and subcritical diffusions. For supercritical diffusions asymptotic forms of distributions are also given in the limit of unbounded exponential growth.

math.PR

Computing expected moments of the Rényi parking problem on the circle

A highly accurate and efficient method to compute the expected values of the count, sum, and squared norm of the sum of the centre vectors of a random maximal sized collection of non-overlapping unit diameter disks touching a fixed unit-diameter disk is presented. This extends earlier work on Rényi's parking problem [Magyar Tud. Akad. Mat. Kutató Int. Közl. 3 (1-2), 1958, pp. 109-127]. Underlying the method is a splitting of the the problem conditional on the value of the first disk. This splitting is proven and then used to derive integral equations for the expectations. These equations take a lower block triangular form. They are solved using substitution and approximation of the integrals to very high accuracy using a polynomial approximation within the blocks.

math.NA

Coalescence and sampling distributions for Feller diffusions

Consider the diffusion process defined by the forward equation $u_t(t, x) = \tfrac{1}{2}\{x u(t, x)\}_{xx} - α\{x u(t, x)\}_{x}$ for $t, x \ge 0$ and $-\infty < α< \infty$, with an initial condition $u(0, x) = δ(x - x_0)$. This equation was introduced and solved by Feller to model the growth of a population of independently reproducing individuals. We explore important coalescent processes related to Feller's solution. For any $α$ and $x_0 > 0$ we calculate the distribution of the random variable $A_n(s; t)$, defined as the finite number of ancestors at a time $s$ in the past of a sample of size $n$ taken from the infinite population of a Feller diffusion at a time $t$ since since its initiation. In a subcritical diffusion we find the distribution of population and sample coalescent trees from time $t$ back, conditional on non-extinction as $t \to \infty$. In a supercritical diffusion we construct a coalescent tree which has a single founder and derive the distribution of coalescent times.

math.PR

The stationary and quasi-stationary properties of neutral multi-type branching process diffusions

The stationary asymptotic properties of the diffusion limit of a multi-type branching process with neutral mutations are studied. For the critical and subcritical processes the interesting limits are those of quasi-stationary distributions conditioned on non-extinction. Pedagogical derivations are given for known results that the limiting distributions for supercritical and critical processes are found to collapse onto rays aligned with stationary eigenvectors of the mutation rate matrix, in agreement with discrete multi-type branching processes. For the sub-critical process the previously unsolved quasi-stationary distribution is obtained to first order in the overall mutation rate, which is assumed to be small. The sampling distribution over allele types for a sample of given finite size is found to agree to first order in mutation rates with the analogous sampling distribution for a Wright-Fisher diffusion with constant population size.

math.PR

Maximum likelihood estimators for scaled mutation rates in an equilibrium mutation-drift model

The stationary sampling distribution of a neutral decoupled Moran or Wright-Fisher diffusion with neutral mutations is known to first order for a general rate matrix with small but otherwise unconstrained mutation rates. Using this distribution as a starting point we derive results for maximum likelihood estimates of scaled mutation rates from site frequency data under three model assumptions: a twelve-parameter general rate matrix, a nine-parameter reversible rate matrix, and a six-parameter strand-symmetric rate matrix. The site frequency spectrum is assumed to be sampled from a fixed size population in equilibrium, and to consist of allele frequency data at a large number of unlinked sites evolving with a common mutation rate matrix without selective bias. We correct an error in a previous treatment of the same problem (Burden and Tang, 2017) affecting the estimators for the general and strand-symmetric rate matrices. The method is applied to a biological dataset consisting of a site frequency spectrum extracted from short autosomal introns in a sample of Drosophila melanogaster individuals.

q-bio.PE

The transition distribution of a sample from a Wright-Fisher diffusion with general small mutation rates

The transition distribution of a sample taken from a Wright-Fisher diffusion with general small mutation rates is found using a coalescent approach. The approximation is equivalent to having at most one mutation in the coalescent tree of the sample up to the most recent common ancestor with additional mutations occurring on the lineage from the most recent common ancestor to the time origin if complete coalescence occurs before the origin. The sampling distribution leads to an approximation for the transition density in the diffusion with small mutation rates. This new solution has interest because the transition density in a Wright-Fisher diffusion with general mutation rates is not known.

q-bio.PE

Coalescence in the diffusion limit of a Bienayme-Galton-Watson branching process

We consider the problem of estimating the elapsed time since the most recent common ancestor of a finite random sample drawn from a population which has evolved through a Bienayme-Galton-Watson branching process. More specifically, we are interested in the diffusion limit appropriate to a supercritical process in the near-critical limit evolving over a large number of time steps. Our approach differs from earlier analyses in that we assume the only known information is the mean and variance of the number of offspring per parent, the observed total population size at the time of sampling, and the size of the sample. We obtain a formula for the probability that a finite random sample of the population is descended from a single ancestor in the initial population, and derive a confidence interval for the initial population size in terms of the final population size and the time since initiating the process. We also determine a joint likelihood surface from which confidence regions can be determined for simultaneously estimating two parameters, (1) the population size at the time of the most recent common ancestor, and (2) the time elapsed since the existence of the most recent common ancestor.

q-bio.PE

The stationary distribution of a sample from the Wright-Fisher diffusion model with general small mutation rates

The stationary distribution of a sample taken from a Wright-Fisher diffusion with general small mutation rates is found using a coalescent approach. The approximation is equivalent to having at most one mutation in the coalescent tree to the first order in the rates. The sample probabilities characterize an approximation for the stationary distribution from the Wright-Fisher diffusion. The approach is different from Burden and Tang (2016,2017) who use a probability flux argument to obtain the same results from a forward diffusion generator equation. The solution has interest because the solution is not known when rates are not small. An analogous solution is found for the configuration of alleles in a general exchangeable binary coalescent tree. In particular an explicit solution is found for a pure birth process tree when individuals reproduce at rate lambda.

q-bio.PE

Stationary distribution of a 2-island 2-allele Wright-Fisher diffusion model with slow mutation and migration rates

The stationary distribution of the diffusion limit of the 2-island, 2-allele Wright-Fisher with small but otherwise arbitrary mutation and migration rates is investigated. Following a method developed by Burden and Tang (2016, 2017) for approximating the forward Kolmogorov equation, the stationary distribution is obtained to leading order as a set of line densities on the edges of the sample space, corresponding to states for which one island is bi-allelic and the other island is non-segregating, and a set of point masses at the corners of the sample space, corresponding to states for which both islands are simultaneously non-segregating. Analytic results for the corner probabilities and line densities are verified independently using the backward generator and for the corner probabilities using the coalescent.

q-bio.PE

Mutation in Populations Governed by a Galton-Watson Branching Process

A population genetics model based on a multitype branching process, or equivalently a Galton-Watson branching process for multiple alleles, is pre- sented. The diffusion limit forward Kolmogorov equation is derived for the case of neutral mutations. The asymptotic stationary solution is obtained and has the property that the extant population partitions into subpopulations whose relative sizes are determined by mutation rates. An approximate time-dependent solution is obtained in the limit of low mutation rates. This solution has the property that the system undergoes a rapid transition from a drift-dominated phase to a mutation-dominated phase in which the distribution collapses onto the asymptotic stationary distribution. The changeover point of the transition is determined by the per-generation growth factor and mutation rate. The approximate solution is confirmed using numerical simulations.

q-bio.PE

Rate Matrix Estimation From Site Frequency Data

A procedure is described for estimating evolutionary rate matrices from observed site frequency data. The procedure assumes (1) that the data are obtained from a constant size population evolving according to a stationary Wright-Fisher model; (2) that the data consist of a multiple alignment of a moderate number of sequenced genomes drawn randomly from the population; and (3) that within the genome a large number of independent, neutral sites evolving with with a common mutation rate matrix can be identified. No restrictions are imposed on the scaled rate matrix other than that the off-diagonal elements are positive and <<1, and that the rows sum to zero. In particular the rate matrix is not assumed to be reversible. The key to the method is an approximate stationary solution to the forward Kolmogorov equation for the multi-allele neutral Wright-Fisher model in the limit of low mutation rates.

q-bio.PE

An approximate stationary solution for multi-allele neutral diffusion with low mutation rates

We address the problem of determining the stationary distribution of the multi-allelic, neutral-evolution Wright-Fisher model in the diffusion limit. A full solution to this problem for an arbitrary K x K mutation rate matrix involves solving for the stationary solution of a forward Kolmogorov equation over a (K - 1)-dimensional simplex, and remains intractable. In most practical situations mutations rates are slow on the scale of the diffusion limit and the solution is heavily concentrated on the corners and edges of the simplex. In this paper we present a practical approximate solution for slow mutation rates in the form of a set of line densities along the edges of the simplex. The method of solution relies on parameterising the general non-reversible rate matrix as the sum of a reversible part and a set of (K - 1)(K - 2)/2 independent terms corresponding to fluxes of probability along closed paths around faces of the simplex. The solution is potentially a first step in estimat- ing non-reversible evolutionary rate matrices from observed allele frequency spectra.

q-bio.PE

Estimation of the methylation pattern distribution from deep sequencing data

Motivation: Bisulphite sequencing enables the detection of cytosine methylation. The sequence of the methylation states of cytosines on any given read forms a methylation pattern that carries substantially more information than merely studying the average methylation level at individual positions. In order to understand better the complexity of DNA methylation landscapes in biological samples, it is important to study the diversity of these methylation patterns. However, the accurate quantification of methylation patterns is subject to sequencing errors and spurious signals due to incomplete bisulphite conversion of cytosines. Results: A statistical model is developed which accounts for the distribution of DNA methylation patterns at any given locus. The model incorporates the effects of sequencing errors and spurious reads, and enables estimation of the true underlying distribution of methylation patterns. Conclusions: Calculation of the estimated distribution over methylation patterns is implemented in the R Bioconductor package MPFE. Source code and documentation of the package are also available for download at http://bioconductor.org/packages/3.0/bioc/html/MPFE.html.

q-bio.GN

An R Implementation of the Polya-Aeppli Distribution

An efficient implementation of the Polya-Aeppli, or geometirc compound Poisson, distribution in the statistical programming language R is presented. The implementation is available as the package polyaAeppli and consists of functions for the mass function, cumulative distribution function, quantile function and random variate generation with those parameters conventionally provided for standard univatiate probability distributions in the stats package in R

stat.CO

Physico-chemical modelling of target depletion during hybridisation on oligonulceotide microarrays

The effect of target molecule depletion from the supernatant solution is incorporated into a physico-chemical model of hybridisation on oligonucleotide microarrays. Two possible regimes are identified: local depletion, in which depletion by a given probe feature only affects that particular probe, and global depletion, in which all features responding to a given target species are affected. Examples are given of two existing spike-in data sets experiencing measurable effects of target depletion. The first of these, from an experiment by Suzuki et al. using custom built arrays with a broad range of probe lengths and mismatch positions, is verified to exhibit local and not global depletion. The second dataset, the well known Affymetrix HGU133a latin square experiment is shown to be very well explained by a global depletion model. It is shown that microarray calibrations relying on Langmuir isotherm models which ignore depletion effects will significantly underestimate specific target concentrations. It is also shown that a combined analysis of perfect match and mismatch probe signals in terms of a simple graphical summary, namely the hook curve method, can discriminate between cases of local and global depletion.

q-bio.GN

Characterising the D2 statistic: word matches in biological sequences

Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between two sequences. Recent advances have been made in the characterisation of this statistic and in the approximation of its distribution. Here, these results are extended to the case of approximate word matches. We compute the exact value of the variance of the D2 statistic for the case of a uniform letter distribution, and introduce a method to provide accurate approximations of the variance in the remaining cases. This enables the distribution of D2 to be approximated for typical situations arising in biological research. We apply these results to the identification of cis-regulatory modules, and show that this method detects such sequences with a high accuracy. The ability to approximate the distribution of D2 for both exact and approximate word matches will enable the use of this statistic in a more precise manner for sequence comparison, database searches, and identification of transcription factor binding sites.

q-bio.QM

Empirical distribution of k-word matches in biological sequences

This study focuses on an alignment-free sequence comparison method: the number of words of length k shared between two sequences, also known as the D_2 statistic. The advantages of the use of this statistic over alignment-based methods are firstly that it does not assume that homologous segments are contiguous, and secondly that the algorithm is computationally extremely fast, the runtime being proportional to the size of the sequence under scrutiny. Existing applications of the D_2 statistic include the clustering of related sequences in large EST databases such as the STACK database. Such applications have typically relied on heuristics without any statistical basis. Rigorous statistical characterisations of the distribution of D_2 have subsequently been undertaken, but have focussed on the distribution's asymptotic behaviour, leaving the distribution of D_2 uncharacterised for most practical cases. The work presented here bridges these two worlds to give usable approximations of the distribution of D_2 for ranges of parameters most frequently encountered in the study of biological sequences.

q-bio.QM