SearcharxivSearch

arXiv · 2609.24816

Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk

Abstract

A well-known phenomenon in statistical analyses of populations of phylogenetic trees in the Billera-Holmes-Vogtmann space is that the topology of the Fréchet mean tree can contain multifurcations (i.e., internal nodes with more than two children), which raises the practical question of whether this reflects a population-level branching structure (hard polytomy) or merely sampling variability in the data (soft polytomy). This is an instance of the more general phenomenon of "stickiness" in non-Euclidean statistics, whereby the sample Fréchet mean in certain non-positively curved stratified spaces becomes permanently trapped in a lower-dimensional stratum. In this work, we identify a particular multidimensional random walk embedded within the Fréchet mean process, and we show that the time at which stickiness occurs is determined by the largest last-passage time above zero of the coordinates of this random walk. Using this representation, we develop a fully nonparametric procedure for estimating the probability that trifurcations in a sample Fréchet mean tree will bifurcate at some future time if more observations are collected. Lastly, we apply our methodology to a problem in phylogenetics where we consider whether an observed trifurcation in the species tree of primates, glires, and tree shrews is genuinely trifurcated at the population level.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Adam Quinn Jaffe. 2026-09-21. Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk. https://arxiv.org/abs/2609.24816

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Phase transitions in microbial lineage trees

Microbial populations exhibit high cell-to-cell variability, which fundamentally shapes population behavior. A striking consequence is the existence of phase transitions, where small genetic or environmental changes trigger abrupt shifts in population dynamics. While biological phase transitions have often been proposed, connecting observed behavior to the underlying physics has remained challenging. We combine population genetics with statistical physics to show how phase transitions arise naturally in microbial populations. We highlight the existence of a first-order transition in a model of bacterial plasmid engineering and find a strict lower bound on the number of plasmids that can be stably maintained in a population.

q-bio.PE

Coexistence coalitions in propagule disperser quasi-communities

Many natural ecosystems harbor large numbers of coexisting species competing for far fewer distinct resources, in apparent defiance of the competitive exclusion principle. Various mechanisms have been proposed to explain this apparent paradox, often pertaining to organisms with a two-stage sessile--propagule life cycle. Here we develop a stochastic model class for such propagule disperser communities that combines competition--colonization trade-offs, spatial heterogeneity, demographic stochasticity, as well as inherited trait variation, and recover several classical models as special or limiting cases. Using bifurcation analysis, we classify equilibrium coalitions near the extinction threshold and give sufficient conditions for their realization by macroscopic equilibria away from the threshold, bypassing the costly numerical computation of the actual equilibrium states. Illustrative examples examine the resulting trait distributions and coalition patterns, demonstrating the interactive effects of different coexistence mechanisms.

q-bio.PE

Demographic inference of pathogen-infected populations from partially observed transmission forests

Genomic epidemiology has several methods for identifying which sampled individuals are linked by close proximity in the transmission chain, whether by grouping them into clusters under a genetic distance threshold or by identifying probable direct transmission pairs. Individuals found to have no sampled neighbours---singletons---are usually set aside. We argue that they are informative. Whether any two sampled individuals prove to be linked depends on the size of the sampling frame and the number of independent lineage introductions into it, and thus the balance of linked and unlinked individuals is itself data about these quantities. We formalise this by treating the intersection of a transmission tree with a sampling frame as a labelled rooted forest, from which a fixed number of nodes are sampled uniformly at random. Using the all-minors matrix-tree theorem, we derive a closed-form expression for the number of labelled rooted $k$-forests on $N$ nodes in which a specified set of nodes is independent. From this we obtain, again in closed form, the probability that a sample of a given size contains no linked individuals at all, and the likelihood of an arbitrary observed configuration of clusters, both when the transmission structure within each is known, and when only the cluster sizes are. The approach is closer to a survey, or to mark-recapture, than to conventional phylodynamic model fitting: the information comes from the linkage structure of a single sample alone with no assumptions regarding pathogen dynamics. We outline three applications: power calculations for prospective studies, a test of the uniform sampling assumption that underlies the reading of clusters as transmission hotspots, and demographic inference itself. We finally set out the assumptions that a more flexible implementation would need to relax.

q-bio.PE