SearcharxivSearch

arXiv subjects

David Bryant

Publications and source records attributed to David Bryant.

At least 19 recordsLinked to original sources

Submodular and strongly submodular functions and diversities

Submodular functions and their close relatives play a key role in combinatorial optimization, decision theory and potential theory. Part of their importance and usefulness stems from the connections with convex functions and polytopes. Here we explore connections between these functions and metric theory, with the bridge provided by diversities, a recently developed generalization of metric spaces to (finite) sets rather than just pairs. Both submodular functions and strongly submodular functions correspond to natural classes of diversities. Submodular diversities, as we define them here, are essentially non-decreasing, intersecting submodular functions which vanish on singletons. We prove new geometric embedding results for these diversities. In particular we show that submodular, strongly submodular, and XOS functions can be represented by the generalized circumradius, a set function in convex analysis equal to the amount a given convex body needs to be stretched to cover a set of points.

math.CO

LvD: A New Algorithm for Computing the Likelihood of a Phylogeny

There are few, if any, algorithms in statistical phylogenetics which are used more heavily than Felsenstein's 1973 pruning method for computing the likelihood of a tree. We present LvD, (Likelihood via Decomposition), an alternative to Felsenstein's algorithm based on a different decomposition of the underlying phylogeny. It works for all standard nucleotide models. The new algorithm allows updates of the likelihood calculation in worst case $O(\log n)$ time with $n$ taxa, as opposed to worst case $O(n)$ time for existing methods. In practice this leads to appreciable improvements in likelihood calculations, the extent of speed-up depending on how balanced or unbalanced the trees are. We explore implications for parallel computing, and show that the approach allows likelihoods to be computed in $O(\log n)$ parallel time per site, compared to (worst case) $O(n)$ time. We implemented and applied the algorithm to large numbers of simulated and empirical data sets and showed that these theoretical advances lead to a significant practical speed-up, although the extent of the improvement depends on how balanced the phylogenies already are.

q-bio.PE

Subtree Distances, Tight Spans and Diversities

Metric embeddings are central to metric theory and its applications. Here we consider embeddings of a different sort: maps from a set to subsets of a metric space so that distances between points are approximated by minimal distances between subsets. Our main result is a characterization of when a set of distances $d(x,y)$ between elements in a set $X$ have a subtree representation, a real tree $T$ and a collection $\{S_x\}_{x \in X}$ of subtrees of~$T$ such that $d(x,y)$ equals the length of the shortest path in~$T$ from a point in $S_x$ to a point in $S_y$ for all $x,y \in X$. The characterization was first established for {\em finite} $X$ by Hirai (2006) using a tight span construction defined for distance spaces, metric spaces without the triangle inequality. To extend Hirai's result beyond finite $X$ we establish fundamental results of tight span theory for general distance spaces, including the surprising observation that the tight span of a distance space is hyperconvex. We apply the results to obtain the first characterization of when a diversity -- a generalization of a metric space which assigns values to all finite subsets of $X$, not just to pairs -- has a tight span which is tree-like.

math.MG

Linear and Sublinear Diversities

Diversities are an extension of the concept of a metric space which assign a non-negative value to every finite set of points, rather than just pairs. A general theory of diversities has been developed which exhibits many deep analogies to metric space theory but also veers off in new directions. Just as many of the most important aspects of metric space theory involve metrics defined on $\mathbb{R}^k$, many applications of diversity theory require a specialized theory for diversities defined on $\mathbb{R}^k$, as we develop here. We focus on two fundamental classes of diversities defined on $\mathbb{R}^k$: those that are Minkowski linear and those that are Minkowski sublinear. Many well-known functions in convex analysis belong to these classes, including diameter, circumradius and mean width. We derive surprising characterizations of these classes, and establish elegant connections between them. Motivated by classical results in metric geometry, and connections with combinatorial optimization, we then examine embeddability of finite diversities into $\mathbb{R}^k$. We prove that a finite diversity can be embedded into a linear diversity exactly when it is of negative type and that it can be embedded into a sublinear diversity exactly when it corresponds to a generalized circumradius.

math.MG

Parsimony and the rank of a flattening matrix

The standard models of sequence evolution on a tree determine probabilities for every character or site pattern. A flattening is an arrangement of these probabilities into a matrix, with rows corresponding to all possible site patterns for one set $A$ of taxa and columns corresponding to all site patterns for another set $B$ of taxa. Flattenings have been used to prove difficult results relating to phylogenetic invariants and consistency and also form the basis of several methods of phylogenetic inference. We prove that the rank of the flattening equals $r^{\ell_T(A|B)}$, where $r$ is the number of states and $\ell_T(A|B)$ is the parsimony length of the binary character separating $A$ and $B$. This result corrects an earlier published formula and opens up new applications for old parsimony theorems. Since completing this work, we have learnt that an equivalent result has been proved much earlier by Casanellas and Fern\'andez-S\'anchez, using a different proof strategy.

q-bio.PE

Diversities and the Generalized Circumradius

The generalized circumradius of a set of points $A \subseteq \mathbb{R}^d$ with respect to a convex body $K$ equals the minimum value of $\lambda \geq 0$ such that $A$ is contained in a translate of $\lambda K$. Each choice of $K$ gives a different function on the set of bounded subsets of $\mathbb{R}^d$; we characterize which functions can arise in this way. Our characterization draws on the theory of diversities, a recently introduced generalization of metrics from functions on pairs to functions on finite subsets. We additionally investigate functions which arise by restricting the generalised circumradius to a finite subset of $\mathbb{R}^d$. We obtain elegant characterizations in the case that $K$ is a simplex or parallelotope.

math.MG

The Geometry of the space of Discrete Coalescent Trees

Computational inference of dated evolutionary histories relies upon various hypotheses about RNA, DNA, and protein sequence mutation rates. Using mutation rates to infer these dated histories is referred to as molecular clock assumption. Coalescent theory is a popular class of evolutionary models that implements the molecular clock hypothesis to facilitate computational inference of dated phylogenies. Cancer and virus evolution are two areas where these methods are particularly important. Methodologically, phylogenetic inference methods require a tree space over which the inference is performed, and geometry of this space plays an important role in statistical and computational aspects of tree inference algorithms. It has recently been shown that molecular clock, and hence coalescent, trees possess a unique geometry, different from that of classical phylogenetic tree spaces which do not model mutation rates. Here we introduce and study a space of discrete coalescent trees, that is, we assume that time is discrete, which is inevitable in many computational formalisations. We establish several geometrical properties of the space and show how these properties impact various algorithms used in phylogenetic analyses. Our tree space is a discretisation of a known time tree space, called t-space, and hence our results can be used to approximate solutions to various open problems in t-space. Our tree space is also a generalisation of another known trees space, called the ranked nearest neighbour interchange space, hence our advances in this paper imply new and generalise existing results about ranked trees.

q-bio.PE

Lattice Diversities

Diversities are a generalization of metric spaces, where instead of the non-negative function being defined on pairs of points, it is defined on arbitrary finite sets of points. Diversities have a well-developed theory. This includes the concept of a diversity tight span that extends the metric tight span in a natural way. Here we explore the generalization of diversities to lattices. Instead of defining diversities on finite subsets of a set we consider diversities defined on members of an arbitrary lattice (with a 0). We show that many of the basic properties of diversities continue to hold. However, the natural map from a lattice diversity to its tight span is not a lattice homomorphism, preventing the development of a complete tight span theory as in the metric and diversity cases.

math.MG

A 1000-fold Acceleration of Hidden Markov Model Fitting using Graphical Processing Units, with application to Nonvolcanic Tremor Classification

Hidden Markov models (HMMs) are general purpose models for time-series data widely used across the sciences because of their flexibility and elegance. However fitting HMMs can often be computationally demanding and time consuming, particularly when the the number of hidden states is large or the Markov chain itself is long. Here we introduce a new Graphical Processing Unit (GPU) based algorithm designed to fit long chain HMMs, applying our approach to an HMM for nonvolcanic tremor events developed by Wang et al.(2018). Even on a modest GPU, our implementation resulted in a 1000-fold increase in speed over the standard single processor algorithm, allowing a full Bayesian inference of uncertainty related to model parameters. Similar improvements would be expected for HMM models given large number of observations and moderate state spaces (<80 states with current hardware). We discuss the model, general GPU architecture and algorithms and report performance of the method on a tremor dataset from the Shikoku region, Japan.

stat.CO

Bayesian inference of species trees using diffusion models

We describe a new and computationally efficient Bayesian methodology for inferring species trees and demographics from unlinked binary markers. Likelihood calculations are carried out using diffusion models of allele frequency dynamics combined with a new algorithm for numerically computing likelihoods of quantitative traits. The diffusion approach allows for analysis of datasets containing hundreds or thousands of individuals. The method, which we call \snapper, has been implemented as part of the Beast2 package. We introduce the models, the efficient algorithms, and report performance of \snapper on simulated data sets and on SNP data from rattlesnakes and freshwater turtles.

q-bio.PE

Fra\"iss\'e Limits for Relational Metric Structures

The general theory developed by Ben Yaacov for metric structures provides Fra\"iss\'e limits which are approximately ultrahomogeneous. We show here that this result can be strengthened in the case of relational metric structures. We give an extra condition that guarantees exact ultrahomogenous limits. The condition is quite general. We apply it to stochastic processes, the class of diversities, and its subclass of $L_1$ diversities.

math.LO

Inner products for Convex Bodies

We define a set inner product to be a function on pairs of convex bodies which is symmetric, Minkowski linear in each dimension, positive definite, and satisfies the natural analogue of the Cauchy-Schwartz inequality (which is not implied by the other conditions). We show that any set inner product can be embedded into an inner product space on the associated support functions, thereby extending fundamental results of Hormander and Radstrom. The set inner product provides a geometry on the space of convex bodies. We explore some of the properties of that geometry, and discuss an application of these ideas to the reconstruction of ancestral ecological niches in evolutionary biology.

math.MG

MAD roots for large trees

The Minimal Ancestral Deviation (MAD) method is a recently introduced procedure for estimating the root of a phylogenetic tree, based only on the shape and branch lengths of the tree. The method is loosely derived from the midpoint rooting method, but, unlike its predecessor, makes use of all pairs of OTUs when positioning the root. In this note we establish properties of this method and then describe a fast and memory efficient algorithm. As a proof of principle, we use our algorithm to determine the MAD roots for simulated phylogenies with up to 100,000 OTUs. The calculations take a few minutes on a standard laptop.

q-bio.PE

An $O(n \log n)$ time Algorithm for computing the Path-length Distance between Trees

Tree comparison metrics have proven to be an invaluable aide in the reconstruction and analysis of phylogenetic (evolutionary) trees. The path-length distance between trees is a particularly attractive measure as it reflects differences in tree shape as well as differences between branch lengths. The distance equals the sum, over all pairs of taxa, of the squared differences between the lengths of the unique path connecting them in each tree. We describe an $O(n \log n)$ time for computing this distance, making extensive use of tree decomposition techniques introduced by Brodal et al. (2004).

cs.DS

Negative type diversities, a multi-dimensional analogue of negative type metrics

Diversities are a generalization of metric spaces in which a non-negative value is assigned to all finite subsets of a set, rather than just to pairs of points. Here we provide an analogue of the theory of negative type metrics for diversities. We introduce negative type diversities, and show that, as in the metric space case, they are a generalization of $L_1$-embeddable diversities. We provide a number of characterizations of negative type diversities, including a geometric characterisation. Much of the recent interest in negative type metrics stems from the connections between metric embeddings and approximation algorithms. We extend some of this work into the diversity setting, showing that lower bounds for embeddings of negative type metrics into $L_1$ can be extended to diversities by using recently established extremal results on hypergraphs.

math.MG

Adaptive Sequential MCMC for Combined State and Parameter Estimation

In the case of a linear state space model, we implement an MCMC sampler with two phases. In the learning phase, a self-tuning sampler is used to learn the parameter mean and covariance structure. In the estimation phase, the parameter mean and covariance structure informs the proposed mechanism and is also used in a delayed-acceptance algorithm. Information on the resulting state of the system is given by a Gaussian mixture. In on-line mode, the algorithm is adaptive and uses a sliding window approach to accelerate sampling speed and to maintain appropriate acceptance rates. We apply the algorithm to joined state and parameter estimation in the case of irregularly sampled GPS time series data.

stat.AP

V-Splines and Bayes Estimate

Smoothing splines can be thought of as the posterior mean of a Gaussian process regression in a certain limit. By constructing a reproducing kernel Hilbert space with an appropriate inner product, the Bayesian form of the V-spline is derived when the penalty term is a fixed constant instead of a function. An extension to the usual generalized cross-validation formula is utilized to find the optimal V-spline parameters.

math.ST

Adaptive Smoothing for Trajectory Reconstruction

Trajectory reconstruction is the process of inferring the path of a moving object between successive observations. In this paper, we propose a smoothing spline -- which we name the V-spline -- that incorporates position and velocity information and a penalty term that controls acceleration. We introduce a particular adaptive V-spline designed to control the impact of irregularly sampled observations and noisy velocity measurements. A cross-validation scheme for estimating the V-spline parameters is given and we detail the performance of the V-spline on four particularly challenging test datasets. Finally, an application of the V-spline to vehicle trajectory reconstruction in two dimensions is given, in which the penalty term is allowed to further depend on known operational characteristics of the vehicle.

stat.ME