SearcharxivSearch

arXiv subjects

Steven N. Evans

Publications and source records attributed to Steven N. Evans.

At least 19 recordsLinked to original sources

Mean-field interacting multi-type birth-death processes with a view to applications in phylodynamics

Multi-type birth-death processes underlie approaches for inferring evolutionary dynamics from phylogenetic trees across biological scales, ranging from deep-time species macroevolution to rapid viral evolution and somatic cellular proliferation. A limitation of current phylogenetic birth-death models is that they require restrictive linearity assumptions that yield tractable message-passing likelihoods, but that also preclude interactions between individuals. Many fundamental evolutionary processes -- such as environmental carrying capacity or frequency-dependent selection -- entail interactions, and may strongly influence the dynamics in some systems. Here, we introduce a multi-type birth-death process in mean-field interaction with an ensemble of replicas of the focal process. We prove that, under quite general conditions, the ensemble's stochastically evolving interaction field converges to a deterministic trajectory in the limit of an infinite ensemble. In this limit, the replicas effectively decouple, and self-consistent interactions appear as nonlinearities in the infinitesimal generator of the focal process. We investigate a special case that is rich enough to model both carrying capacity and frequency-dependent selection while yielding tractable message-passing likelihoods in the context of a phylogenetic birth-death model.

math.PR

Progress on Constructing Phylogenetic Networks for Languages

In 2006, Warnow, Evans, Ringe, and Nakhleh proposed a stochastic model (hereafter, the WERN 2006 model) of multi-state linguistic character evolution that allowed for homoplasy and borrowing. They proved that if there is no borrowing between languages and homoplastic states are known in advance, then the phylogenetic tree of a set of languages is statistically identifiable under this model, and they presented statistically consistent methods for estimating these phylogenetic trees. However, they left open the question of whether a phylogenetic network -- which would explicitly model borrowing between languages that are in contact -- can be estimated under the model of character evolution. Here, we establish that under some mild additional constraints on the WERN 2006 model, the phylogenetic network topology is statistically identifiable, and we present algorithms to infer the phylogenetic network. We discuss the ramifications for linguistic phylogenetic network estimation in practice, and suggest directions for future research.

q-bio.PE

Limit Theorems for Fréchet Mean Sets

For $1\le p \le \infty$, the Fréchet $p$-mean of a probability measure on a metric space is an important notion of central tendency that generalizes the usual notions in the real line of mean ($p=2$) and median ($p=1$). In this work we prove a collection of limit theorems for Fréchet means and related objects, which, in general, constitute a sequence of random closed sets. On the one hand, we show that many limit theorems (a strong law of large numbers, an ergodic theorem, and a large deviations principle) can be simply descended from analogous theorems on the space of probability measures via purely topological considerations. On the other hand, we provide the first sufficient conditions for the strong law of large numbers to hold in a $T_2$ topology (in particular, the Fell topology), and we show that this condition is necessary in some special cases. We also discuss statistical and computational implications of the results herein.

math.PR

Virtual Markov chains

We introduce the space of virtual Markov chains (VMCs) as a projective limit of the spaces of all finite state space Markov chains (MCs), in the same way that the space of virtual permutations is the projective limit of the spaces of all permutations of finite sets. We introduce the notions of virtual initial distribution (VID) and a virtual transition matrix (VTM), and we show that the law of any VMC is uniquely characterized by a pair of a VID and VTM which have to satisfy a certain compatibility condition. Lastly, we study various properties of compact convex sets associated to the theory of VMCs, including that the Birkhoff-von Neumann theorem fails in the virtual setting.

math.PR

Two continua of embedded regenerative sets

Given a two-sided real-valued Lévy process $(X_t)_{t \in \mathbb{R}}$, define processes $(L_t)_{t \in \mathbb{R}}$ and $(M_t)_{t \in \mathbb{R}}$ by $L_t := \sup\{h \in \mathbb{R} : h - α(t-s) \le X_s \text{ for all } s \le t\} = \inf\{X_s + α(t-s) : s \le t\}$, $t \in \mathbb{R}$, and $M_t := \sup \{ h \in \mathbb{R} : h - α|t-s| \leq X_s \text{ for all } s \in \mathbb{R} \} = \inf \{X_s + α|t-s| : s \in \mathbb{R}\}$, $t \in \mathbb{R}$. The corresponding contact sets are the random sets $\mathcal{H}_α:= \{ t \in \mathbb{R} : X_{t}\wedge X_{t-} = L_t\}$ and $\mathcal{Z}_α:= \{ t \in \mathbb{R} : X_{t}\wedge X_{t-} = M_t\}$. For a fixed $α>\mathbb{E}[X_1]$ (resp. $α>|\mathbb{E}[X_1]|$) the set $\mathcal{H}_α$ (resp. $\mathcal{Z}_α$) is non-empty, closed, unbounded above and below, stationary, and regenerative. The collections $(\mathcal{H}_α)_{α> \mathbb{E}[X_1]}$ and $(\mathcal{Z}_α)_{α> |\mathbb{E}[X_1]|}$ are increasing in $α$ and the regeneration property is compatible with these inclusions in that each family is a continuum of embedded regenerative sets in the sense of Bertoin. We show that $(\sup\{t < 0 : t \in \mathcal{H}_α\})_{α> \mathbb{E}[X_1]}$ is a càdlàg, nondecreasing, pure jump process with independent increments and determine the intensity measure of the associated Poisson process of jumps. We obtain a similar result for $(\sup\{t < 0 : t \in \mathcal{Z}_α\})_{α> |β|}$ when $(X_t)_{t \in \mathbb{R}}$ is a (two-sided) Brownian motion with drift $β$.

math.PR

Doob--Martin boundary of Rémy's tree growth chain

Rémy's algorithm is a Markov chain that iteratively generates a sequence of random trees in such a way that the $n^{\mathrm{th}}$ tree is uniformly distributed over the set of rooted, planar, binary trees with $2n+1$ vertices. We obtain a concrete characterization of the Doob--Martin boundary of this transient Markov chain and thereby delineate all the ways in which, loosely speaking, this process can be conditioned to "go to infinity" at large times. A (deterministic) sequence of finite rooted, planar, binary trees converges to a point in the boundary if for each $m$ the random rooted, planar, binary tree spanned by $m+1$ leaves chosen uniformly at random from the $n^{\mathrm{th}}$ tree in the sequence converges in distribution as $n$ tends to infinity -- a notion of convergence that is analogous to one that appears in the recently developed theory of graph limits. We show that a point in the Doob--Martin boundary may be identified with the following ensemble of objects: a complete separable $\mathbb{R}$-tree that is rooted and binary in a suitable sense, a diffuse probability measure on the $\mathbb{R}$-tree that allows us to make sense of sampling points from it, and a kernel on the $\mathbb{R}$-tree that describes the probability that the first of a given pair of points is below and to the left of their most recent common ancestor while the second is below and to the right. The Doob--Martin boundary corresponds bijectively to the set of extreme points of the closed convex set of normalized nonnegative harmonic functions, in other words, the minimal and full Doob--Martin boundaries coincide. These results are in the spirit of the identification of graphons as limit objects in the theory of graph limits.

math.PR

PATRICIA bridges

A radix sort tree arises when storing distinct infinite binary words in the leaves of a binary tree such that for any two words their common prefixes coincide with the common prefixes of the corresponding two leaves. If one deletes the out-degree $1$ vertices in the radix sort tree and "closes up the gaps", then the resulting PATRICIA tree maintains all the information that is necessary for sorting the infinite words into lexicographic order. We investigate the PATRICIA chains -- the tree-valued Markov chains that arise when successively building the PATRICIA trees for the collection of infinite binary words $Z_1,\ldots, Z_n$, $n=1,2,\ldots$, where the source words $Z_1, Z_2,\ldots$ are independent and have a common diffuse distribution on $\{0,1\}^\infty$. It turns out that the PATRICIA chains share a common collection of backward transition probabilities and that these are the same as those of a chain introduced by Rémy for successively generating uniform random binary trees with larger and larger numbers of leaves. This means that the infinite bridges of any PATRICIA chain (that is, the chains obtained by conditioning a PATRICIA chain on its remote future) coincide with the infinite bridges of the Rémy chain. The infinite bridges of the Rémy chain are characterized concretely in Evans, Grübel, and Wakolbinger 2017 and we recall that characterization here while adding some details and clarifications.

math.PR

Rotatable random sequences in local fields

An infinite sequence of real random variables $(ξ_1, ξ_2, \dots)$ is said to be rotatable if every finite subsequence $(ξ_1, \dots, ξ_n)$ has a spherically symmetric distribution. A celebrated theorem of Freedman states that $(ξ_1, ξ_2, \dots)$ is rotatable if and only if $ξ_j = τη_j$ for all $j$, where $(η_1, η_2, \dots)$ is a sequence of independent standard Gaussian random variables and $τ$ is an independent nonnegative random variable. Freedman's theorem is equivalent to a classical result of Schoenberg which says that a continuous function $ϕ: \mathbb{R}_+ \to \mathbb{C}$ with $ϕ(0) = 1$ is completely monotone if and only if $ϕ_n: \mathbb{R}^n \to \mathbb{R}$ given by $ϕ_n(x_1, \ldots, x_n) = ϕ(x_1^2 + \cdots + x_n^2)$ is nonnegative definite for all $n \in \mathbb{N}$. We establish the analogue of Freedman's theorem for sequences of random variables taking values in local fields using probabilistic methods and then use it to establish a local field analogue of Schoenberg's result. Along the way, we obtain a local field counterpart of an observation variously attributed to Maxwell, Poincaré, and Borel which says that if $(ζ_1, \ldots, ζ_n)$ is uniformly distributed on the sphere of radius $\sqrt{n}$ in $\mathbb{R}^n$, then, for fixed $k \in \mathbb{N}$, the distribution of $(ζ_1, \ldots, ζ_k)$ converges to that of a vector of $k$ independent standard Gaussian random variables as $n \to \infty$.

math.PR

Excursions away from the Lipschitz minorant of a Lévy process

For $α>0$, the $α$-Lipschitz minorant of a function $f : \mathbb{R} \rightarrow \mathbb{R}$ is the greatest function $m : \mathbb{R} \rightarrow \mathbb{R}$ such that $m \leq f$ and $\vert m(s) - m(t) \vert \leq α\vert s-t \vert$ for all $s,t \in \mathbb{R}$, should such a function exist. If $X=(X_t)_{t \in \mathbb{R}}$ is a real-valued Lévy process that is not a pure linear drift with slope $\pm α$, then the sample paths of $X$ have an $α$-Lipschitz minorant almost surely if and only if $\mathbb{E}[\vert X_1 \vert]< \infty$ and $\vert \mathbb{E}[X_1]\vert < α$. Denoting the minorant by $M$, we consider the contact set $\mathcal{Z}:=\{ t \in \mathbb{R} : M_t = X_t \wedge X_{t-}\}$, which, since it is regenerative and stationary, has the distribution of the closed range of some subordinator "made stationary" in a suitable sense. We provide a description of the excursions of the Lévy process away from its contact set similar to the one presented in Itô excursion theory. We study the distribution of the excursion on the special interval straddling zero. We also give an explicit path decomposition of the other "generic" excursions in the case of Brownian motion with drift $β$ with $\vert β\vert < α$. Finally, we investigate the progressive enlargement of the Brownian filtration by the random time that is the first point of the contact set after zero.

math.PR

Representing smooth functions as compositions of near-identity functions with implications for deep network optimization

We show that any smooth bi-Lipschitz $h$ can be represented exactly as a composition $h_m \circ ... \circ h_1$ of functions $h_1,...,h_m$ that are close to the identity in the sense that each $\left(h_i-\mathrm{Id}\right)$ is Lipschitz, and the Lipschitz constant decreases inversely with the number $m$ of functions composed. This implies that $h$ can be represented to any accuracy by a deep residual network whose nonlinear layers compute functions with a small Lipschitz constant. Next, we consider nonlinear regression with a composition of near-identity nonlinear maps. We show that, regarding Fréchet derivatives with respect to the $h_1,...,h_m$, any critical point of a quadratic criterion in this near-identity region must be a global minimizer. In contrast, if we consider derivatives with respect to parameters of a fixed-size residual network with sigmoid activation functions, we show that there are near-identity critical points that are suboptimal, even in the realizable case. Informally, this means that functional gradient methods for residual networks cannot get stuck at suboptimal critical points corresponding to near-identity layers, whereas parametric gradient methods for sigmoidal residual networks suffer from suboptimal critical points in the near-identity region.

cs.LG

The spans in Brownian motion

For $d \in \{1,2,3\}$, let $(B^d_t;~ t \geq 0)$ be a $d$-dimensional standard Brownian motion. We study the $d$-Brownian span set $Span(d):=\{t-s;~ B^d_s=B^d_t~\mbox{for some}~0 \leq s \leq t\}$. We prove that almost surely the random set $Span(d)$ is $σ$-compact and dense in $\mathbb{R}_{+}$. In addition, we show that $Span(1)=\mathbb{R}_{+}$ almost surely; the Lebesgue measure of $Span(2)$ is $0$ almost surely and its Hausdorff dimension is $1$ almost surely; and the Hausdorff dimension of $Span(3)$ is $\frac{1}{2}$ almost surely. We also list a number of conjectures and open problems.

math.PR

Doob-Martin compactification of a Markov chain for growing random words sequentially

We consider a Markov chain that iteratively generates a sequence of random finite words in such a way that the $n^{\mathrm{th}}$ word is uniformly distributed over the set of words of length $2n$ in which $n$ letters are $a$ and $n$ letters are $b$: at each step an $a$ and a $b$ are shuffled in uniformly at random among the letters of the current word. We obtain a concrete characterization of the Doob-Martin boundary of this Markov chain. Writing $N(u)$ for the number of letters $a$ (equivalently, $b$) in the finite word $u$, we show that a sequence $(u_n)_{n \in \mathbb{N}}$ of finite words converges to a point in the boundary if, for an arbitrary word $v$, there is convergence as $n$ tends to infinity of the probability that the selection of $N(v)$ letters $a$ and $N(v)$ letters $b$ uniformly at random from $u_n$ and maintaining their relative order results in $v$. We exhibit a bijective correspondence between the points in the boundary and ergodic random total orders on the set $\{a_1, b_1, a_2, b_2, \ldots \}$ that have distributions which are separately invariant under finite permutations of the indices of the $a'$s and those of the $b'$s. We establish a further bijective correspondence between the set of such random total orders and the set of pairs $(μ,ν)$ of diffuse probability measures on $[0,1]$ such that $\frac{1}{2}(μ+ν)$ is Lebesgue measure: the restriction of the random total order to $\{a_1, b_1, \ldots, a_n, b_n\}$ is obtained by taking $X_1, \ldots, X_n$ (resp. $Y_1, \ldots, Y_n$) i.i.d. with common distribution $μ$ (resp. $ν$), letting $(Z_1, \ldots, Z_{2n})$ be $\{X_1, Y_1, \ldots, X_n, Y_n\}$ in increasing order, and declaring that the $k^{\mathrm{th}}$ smallest element in the restricted total order is $a_i$ (resp. $b_j$) if $Z_k = X_i$ (resp. $Z_k = Y_j$).

math.PR

Leading the field: Fortune favors the bold in Thurstonian choice models

Schools with the highest average student performance are often the smallest schools; localities with the highest rates of some cancers are frequently small and the effects observed in clinical trials are likely to be largest for the smallest numbers of subjects. Informal explanations of this "small-schools phenomenon" point to the fact that the sample means of smaller samples have higher variances. But this cannot be a complete explanation: If we draw two samples from a diffuse distribution that is symmetric about some point, then the chance that the smaller sample has larger mean is 50\%. A particular consequence of results proved below is that if one draws three or more samples of different sizes from the same normal distribution, then the sample mean of the smallest sample is most likely to be highest, the sample mean of the second smallest sample is second most likely to be highest, and so on; this is true even though for any pair of samples, each one of the pair is equally likely to have the larger sample mean. Our conclusions are relevant to certain stochastic choice models including the following generalization of Thurstone's Law of Comparative Judgment. There are $n$ items. Item $i$ is preferred to item $j$ if $Z_i < Z_j$, where $Z$ is a random $n$-vector of preference scores. Suppose $\mathbb{P}\{Z_i = Z_j\} = 0$ for $i \ne j$, so there are no ties. Item $k$ is the favorite if $Z_k < \min_{i\ne k} Z_i$. Let $p_i$ denote the chance that item $i$ is the favorite. We characterize a large class of distributions for $Z$ for which $p_1 > p_2 > \cdots > p_n$. Our results are most surprising when $\mathbb{P}\{Z_i < Z_j\} = \mathbb{P}\{Z_i > Z_j\} = \frac{1}{2}$ for $i \ne j$, so neither of any two items is likely to be preferred over the other in a pairwise comparison.

math.PR

Markov processes conditioned on their location at large exponential times

Suppose that $(X_t)_{t \ge 0}$ is a one-dimensional Brownian motion with negative drift $-μ$. It is possible to make sense of conditioning this process to be in the state $0$ at an independent exponential random time and if we kill the conditioned process at the exponential time the resulting process is Markov. If we let the rate parameter of the random time go to $0$, then the limit of the killed Markov process evolves like $X$ conditioned to hit $0$, after which time it behaves as $X$ killed at the last time $X$ visits $0$. Equivalently, the limit process has the dynamics of the killed "bang--bang" Brownian motion that evolves like Brownian motion with positive drift $+μ$ when it is negative, like Brownian motion with negative drift $-μ$ when it is positive, and is killed according to the local time spent at $0$. An extension of this result holds in great generality for Borel right processes conditioned to be in some state $a$ at an exponential random time, at which time they are killed. Our proofs involve understanding the Campbell measures associated with local times, the use of excursion theory, and the development of a suitable analogue of the "bang--bang" construction for general Markov processes. As examples, we consider the special case when the transient Borel right process is a one-dimensional diffusion. Characterizing the limiting conditioned and killed process via its infinitesimal generator leads to an investigation of the $h$-transforms of transient one-dimensional diffusion processes that goes beyond what is known and is of independent interest.

math.PR

Radix sort trees in the large

The trie-based radix sort algorithm stores pairwise different infinite binary strings in the leaves of a binary tree in a way that the Ulam-Harris coding of each leaf equals a prefix (that is, an initial segment) of the corresponding string, with the prefixes being of minimal length so that they are pairwise different. We investigate the {\em radix sort tree chains} -- the tree-valued Markov chains that arise when successively storing infinite binary strings $Z_1,\ldots, Z_n$, $n=1,2,\ldots$ according to the trie-based radix sort algorithm, where the source strings $Z_1, Z_2,\ldots$ are independent and identically distributed. We establish a bijective correspondence between the full Doob--Martin boundary of the radix sort tree chain with a {\em symmetric Bernoulli source} (that is, each $Z_k$ is a fair coin-tossing sequence) and the family of radix sort tree chains for which the common distribution of the $Z_k$ is a diffuse probability measure on $\{0,1\}^\infty$. In essence, our result characterizes all the ways that it is possible to condition such a chain of radix sort trees consistently on its behavior "in the large".

math.PR

Bayesian inference of natural selection from allele frequency time series

The advent of accessible ancient DNA technology now allows the direct ascertainment of allele frequencies in ancestral populations, thereby enabling the use of allele frequency time series to detect and estimate natural selection. Such direct observations of allele frequency dynamics are expected to be more powerful than inferences made using patterns of linked neutral variation obtained from modern individuals. We develop a Bayesian method to make use of allele frequency time series data and infer the parameters of general diploid selection, along with allele age, in non-equilibrium populations. We introduce a novel path augmentation approach, in which we use Markov chain Monte Carlo to integrate over the space of allele frequency trajectories consistent with the observed data. Using simulations, we show that this approach has good power to estimate selection coefficients and allele age. Moreover, when applying our approach to data on horse coat color, we find that ignoring a relevant demographic history can significantly bias the results of inference. Our approach is made available in a C++ software package.

q-bio.PE

Polar decomposition of scale-homogeneous measures with application to Lévy measures of strictly stable laws

A scaling on some space is a measurable action of the group of positive real numbers. A measure on a measurable space equipped with a scaling is said to be $α$-homogeneous for some nonzero real number $α$ if the mass of any measurable set scaled by any factor $t > 0$ is the multiple $t^{-α}$ of the set's original mass. It is shown rather generally that given an $α$-homogeneous measure on a measurable space there is a measurable bijection between the space and the Cartesian product of a subset of the space and the positive real numbers (that is, a "system of polar coordinates") such that the push-forward of the $α$-homogeneous measure by this bijection is the product of a probability measure on the first component (that is, on the "angular" component) and an $α$-homogeneous measure on the positive half-line (that is, on the "radial" component). This result is applied to the intensity measures of Poisson processes that arise in Lévy-Khinchin-Itô-like representations of infinitely divisible random elements. It is established that if a strictly stable random element in a convex cone admits a series representation as the sum of points of a Poisson process, then it necessarily has a LePage representation as the sum of i.i.d. random elements of the cone scaled by the successive points of an independent unit intensity Poisson process on the positive half-line each raised to the power $-\frac{1}α$.

math.PR

When do skew-products exist?

The classical skew-product decomposition of planar Brownian motion represents the process in polar coordinates as an autonomously Markovian radial part and an angular part that is an independent Brownian motion on the unit circle time-changed according to the radial part. Theorem 4 of Liao (2009) gives a broad generalization of this fact to a setting where there is a diffusion on a manifold $X$ with a distribution that is equivariant under the smooth action of a Lie group $K$. Under appropriate conditions, there is a decomposition into an autonomously Markovian "radial" part that lives on the space of orbits of $K$ and an "angular" part that is an independent Brownian motion on the homogeneous space $K/M$, where $M$ is the isotropy subgroup of a point of $x$, that is time-changed with a time-change that is adapted to the filtration of the radial part. We present two apparent counterexamples to Theorem 4 of Liao (2009). In the first counterexample the angular part is not a time-change of any Brownian motion on $K/M$, whereas in the second counterexample the angular part is the time-change of a Brownian motion on $K/M$ but this Brownian motion is not independent of the radial part. In both of these examples $K/M$ has dimension $1$. The statement and proof of Theorem 4 from Liao (2009) remain valid when $K/M$ has dimension greater than $1$. Our examples raise the question of what conditions lead to the usual sort of skew-product decomposition when $K/M$ has dimension $1$ and what conditions lead to there being no decomposition at all or one in which the angular part is a time-changed Brownian motion but this Brownian motion is not independent of the radial part.

math.PR