Searcharxiv⌕ Search

arXiv subjects

David Bryant

Publications and source records attributed to David Bryant.

35 records · Page 2Linked to original sources

Open Problem Statement: Minimal Distortion Embeddings of Diversities in $\ell_1$

We state an open problem in the theory of diversities: what is the worst case minimal distortion embedding of a diversity on $n$ points in $\ell_1$. This problem is the diversity analogue of a famous problem in metric geometry: what is the worst case minimal distortion embedding of an $n$-point metric space in $\ell_1$. We explain the problem, state some special classes of diversities for which the answer is known, and show why the standard techniques from the metric space case do not work. We then outline some possible lines of attack for the problem that are not yet fully explored.

math.MG↗

A Universal Separable Diversity

The Urysohn space is a separable complete metric space with two fundamental properties: (a) universality: every separable metric space can be isometrically embedded in it; (b) ultrahomogeneity: every finite isometry between two finite subspaces can be extended to an auto-isometry of the whole space. The Urysohn space is uniquely determined up to isometry within separable metric spaces by these two properties. We introduce an analogue of the Urysohn space for diversities, a recently developed variant of the concept of a metric space. In a diversity any finite set of points is assigned a non-negative value, extending the notion of a metric which only applies to unordered pairs of points. We construct the unique separable complete diversity that it is ultrahomogeneous and universal with respect to separable diversities.

math.MG↗

Can we 'future-proof' consensus trees?

Consensus methods are widely used for combining phylogenetic trees into a single estimate of the evolutionary tree for a group of species. As more taxa are added, the new source trees may begin to tell a different evolutionary story when restricted to the original set of taxa. However, if the new trees, restricted to the original set of taxa, were to agree exactly with the earlier trees, then we might hope that their consensus would either agree with or resolve the original consensus tree. In this paper, we ask under what conditions consensus methods exist that are 'future proof' in this sense. While we show that some methods (e.g. Adams consensus) have this property for specific types of input, we also establish a rather surprising `no-go' theorem: there is no 'reasonable' consensus method that satisfies the future-proofing property in general. We then investigate a second notion of 'future proofing' for consensus methods, in which trees (rather than taxa) are added, and establish some positive and negative results. We end with some questions for future work.

q-bio.PE↗

Constant distortion embeddings of Symmetric Diversities

Diversities are like metric spaces, except that every finite subset, instead of just every pair of points, is assigned a value. Just as there is a theory of minimal distortion embeddings of finite metric spaces into $L_1$, there is a similar, yet undeveloped, theory for embedding finite diversities into the diversity analogue of $L_1$ spaces. In the metric case, it is well known that an $n$-point metric space can be embedded into $L_1$ with $\mathcal{O}(\log n)$ distortion. For diversities, the optimal distortion is unknown. Here, we establish the surprising result that symmetric diversities, those in which the diversity (value) assigned to a set depends only on its cardinality, can be embedded in $L_1$ with constant distortion.

math.MG↗

Efficient recycled algorithms for quantitative trait models on phylogenies

We present an efficient and flexible method for computing likelihoods of phenotypic traits on a phylogeny. The method does not resort to Monte-Carlo computation but instead blends Felsenstein's discrete character pruning algorithm with methods for numerical quadrature. It is not limited to Gaussian models and adapts readily to model uncertainty in the observed trait values. We demonstrate the framework by developing efficient algorithms for likelihood calculation and ancestral state reconstruction under Wright's threshold model, applying our methods to a dataset of trait data for extrafloral nectaries (EFNs) across a phylogeny of 839 Labales species.

q-bio.PE↗

When can splits be drawn in the plane?

Split networks are a popular tool for the analysis and visualization of complex evolutionary histories. Every collection of splits (bipartitions) of a finite set can be represented by a split network. Here we characterize which collection of splits can be represented using a planar split network. Our main theorem links these collections of splits with oriented matroids and arrangements of lines separating points in the plane. As a consequence of our main theorem, we establish a particularly simple characterization of maximal collections of these splits.

q-bio.PE↗

Diversities and the Geometry of Hypergraphs

The embedding of finite metrics in $\ell_1$ has become a fundamental tool for both combinatorial optimization and large-scale data analysis. One important application is to network flow problems in which there is close relation between max-flow min-cut theorems and the minimal distortion embeddings of metrics into $\ell_1$. Here we show that this theory can be generalized considerably to encompass Steiner tree packing problems in both graphs and hypergraphs. Instead of the theory of $\ell_1$ metrics and minimal distortion embeddings, the parallel is the theory of diversities recently introduced by Bryant and Tupper, and the corresponding theory of $\ell_1$ diversities and embeddings which we develop here.

math.MG↗

Parsimony via concensus

The parsimony score of a character on a tree equals the number of state changes required to fit that character onto the tree. We show that for unordered, reversible characters this score equals the number of tree rearrangements required to fit the tree onto the character. We discuss implications of this connection for the debate over the use of consensus trees or total evidence, and show how it provides a link between incongruence of characters and recombination.

q-bio.PE↗

Hyperconvexity and Tight Span Theory for Diversities

The tight span, or injective envelope, is an elegant and useful construction that takes a metric space and returns the smallest hyperconvex space into which it can be embedded. The concept has stimulated a large body of theory and has applications to metric classification and data visualisation. Here we introduce a generalisation of metrics, called diversities, and demonstrate that the rich theory associated to metric tight spans and hyperconvexity extends to a seemingly richer theory of diversity tight spans and hyperconvexity.

math.MG↗

Parameter Exploration in Simulation Experiments: A Bayesian Framework

Simulations often involve the use of model parameters which are unknown or uncertain. For this reason, simulation experiments are often repeated for multiple combinations of parameter values, often iterating through parameter values lying on a fixed grid. However, the use of a discrete grid places limits on the dimension of the parameter space and creates the potential to miss important parameter combinations which fall in the gaps between grid points. Here we draw parallels with strategies for numerical integration and describe a Markov chain Monte-Carlo strategy for exploring parameter values. We illustrate the approach using examples from phylogenetics, archaeology, and epidemiology.

stat.CO↗

Exact coalescent likelihoods for unlinked markers in finite-sites mutation models

We derive exact formulae for the allele frequency spectrum under the coalescent with mutation, conditioned on allele counts at some fixed time in the past. We consider unlinked biallelic markers mutating according to a finite sites, or infinite sites, model. This work extends the coalescent theory of unlinked biallelic markers, enabling fast computations of allele frequency spectra in multiple populations. Our results have applications to demographic inference, species tree inference, and the analysis of genetic variation in closely related species more generally.

q-bio.PE↗

Inferring Species Trees Directly from Biallelic Genetic Markers: Bypassing Gene Trees in a Full Coalescent Analysis

The multi-species coalescent provides an elegant theoretical framework for estimating species trees and species demographics from genetic markers. Practical applications of the multi-species coalescent model are, however, limited by the need to integrate or sample over all gene trees possible for each genetic marker. Here we describe a polynomial-time algorithm that computes the likelihood of a species tree directly from the markers under a finite-sites model of mutation, effectively integrating over all possible gene trees. The method applies to independent (unlinked) biallelic markers such as well-spaced single nucleotide polymorphisms (SNPs), and we have implemented it in SNAPP, a Markov chain Monte-Carlo sampler for inferring species trees, divergence dates, and population sizes. We report results from simulation experiments and from an analysis of 1997 amplified fragment length polymorphism (AFLP) loci in 69 individuals sampled from six species of {\em Ourisia} (New Zealand native foxglove).

q-bio.PE↗

'Bureaucratic' set systems, and their role in phylogenetics

We say that a collection $\Cc$ of subsets of $X$ is {\em bureaucratic} if every maximal hierarchy on $X$ contained in $\Cc$ is also maximum. We characterise bureaucratic set systems and show how they arise in phylogenetics. This framework has several useful algorithmic consequences: we generalize some earlier results and derive a polynomial-time algorithm for a parsimony problem arising in phylogenetic networks.

q-bio.PE↗

The link between segregation and phylogenetic diversity

We derive an invertible transform linking two widely used measures of species diversity: phylogenetic diversity and the expected proportions of segregating (non-constant) sites. We assume a bi-allelic, symmetric, finite site model of substitution. Like the Hadamard transform of Hendy and Penny, the transform can be expressed completely independent of the underlying phylogeny. Our results bridge work on diversity from two quite distinct scientific communities.

q-bio.PE↗

Computing the Distribution of a Tree Metric

The Robinson-Foulds (RF) distance is by far the most widely used measure of dissimilarity between trees. Although the distribution of these distances has been investigated for twenty years, an algorithm that is explicitly polynomial time has yet to be described for computing this distribution (which is also the distribution of trees around a given tree under the popular Robinson-Foulds metric). In this paper we derive a polynomial-time algorithm for this distribution. We show how the distribution can be approximated by a Poisson distribution determined by the proportion of leaves that lie in `cherries' of the given tree. We also describe how our results can be used to derive normalization constants that are required in a recently-proposed maximum likelihood approach to supertree construction.

q-bio.PE↗

Hadamard Phylogenetic Methods and the n-taxon process

The Hadamard transform of \cite{Hendy89a, Hendy89} provides a way to work with stochastic models for sequence evolution without having to deal with the complications of tree space and the graphical structure of trees. Here we demonstrate that the transform can be expressed in terms of the familiar $\bP[\bt] = e^{\bQ[\bt]}$ formula for Markov chains. The key idea is to study the evolution of vectors of states, one vector entry for each taxa; we call this the $n$-taxon process. We derive transition probabilities for the process. Significantly, the findings show that tree-based models are indeed in the family of (multi-variate) exponential distributions.

q-bio.PE↗

Continuous and Tractable models for the Variation of Evolutionary Rates

We propose a continuous model for evolutionary rate variation across sites and over the tree and derive exact transition probabilities under this model. Changes in rate are modelled using the CIR process, a diffusion widely used in financial applications. The model directly extends the standard gamma distributed rates across site model, with one additional parameter governing changes in rate down the tree. The parameters of the model can be estimated directly from two well-known statistics: the index of dispersion and the gamma shape parameter of the rates across sites model. The CIR model can be readily incorporated into probabilistic models for sequence evolution. We provide here an exact formula for the likelihood of a three taxa tree. Larger trees can be evaluated using Monte-Carlo methods.

math.PR↗