Searcharxiv⌕ Search

arXiv subjects

Katharina T. Huber

Publications and source records attributed to Katharina T. Huber.

At least 19 recordsLinked to original sources

Characterization of tree-child networks in terms of mu vectors

We characterize tree-child phylogenetic networks in terms of their mu-representations. First, we give a structural characterization of tree-child networks by means of ordered tree-path decompositions. We then translate this decomposition into a set of purely vectorial conditions on finite subsets M in N^n. We prove that such a set M is the mu-representation of a tree-child phylogenetic network if and only if it is tree-child mu-compatible. This provides a feasibility criterion for tree-child mu-representations which can be used as a basis for reconstruction and further algorithmic applications. Note that this paper presents results arising from ongoing research on tree-child networks and that the results will be further developed and placed into proper context in subsequent versions.

math.CO↗

Arboreal Networks and Ultrametrics

Ultametrics are an important class of distances used in applications such as phylogenetics, clustering and classification theory. Ultrametrics are essentially distances that can be represented by an edge-weighted rooted tree so that all of the distances in the tree from the root to any leaf of the tree are equal. In this paper, we introduce a generalization of ultrametrics called arboreal ultrametrics which have applications in phylogenetics and also arise in the theory of distance-hereditary graphs. These are partial distances, that is distances that are not necessarily defined for every pair of elements in the groundset, that can be represented by an ultrametric arboreal network, that is, an edge-weighted rooted network whose underlying graph is a tree. As with ultrametrics all of the distances in the ultrametric arboreal network from any root to any leaf below it are are equal but, in contrast, the network may have more than one root. In our two main results we characterize when a partial distance is an arboreal ultrametric as well as proving that, somewhat surprisingly, given any unrooted edge-weighted phylogenetic tree there is a necessarily unique way to insert roots into this tree so as to obtain an arboreal ultrametric.

math.CO↗

Efficient Reconstruction of Arboreal Networks

Arboreal networks are multi-rooted phylogenetic networks whose underlying graph is a tree. We give an encoding of stack-free arboreal networks in terms of triplets and the novel concept of a duet. This yields a polynomial time algorithm to construct these networks from complete triplet and duet systems. The classification results show correctness and lead to a natural metric on these multi-rooted networks.

cs.DM↗

Characterizing semi-directed phylogenetic networks and their multi-rootable variants

In evolutionary biology, phylogenetic networks are graphs that provide a flexible framework for representing complex evolutionary histories that involve reticulate evolutionary events. Recently phylogenetic studies have started to focus on a special class of such networks called semi-directed networks. These graphs are defined as mixed graphs that can be obtained by de-orienting some of the arcs in some rooted phylogenetic network, that is, a directed acyclic graph whose leaves correspond to a collection of species and that has a single source or root vertex. However, this definition of semi-directed networks is implicit in nature since it is not clear when a mixed-graph enjoys this property or not. In this paper, we introduce novel, explicit mathematical characterizations of semi-directed networks, and also multi-semi-directed networks, that is, mixed graphs that can be obtained from directed phylogenetic networks that may have more than one root. In addition, through extending foundational tools from the theory of rooted networks into the semi-directed setting - such as cherry picking sequences, omnians, and path partitions - we characterize when a (multi-)semi-directed network can be obtained by de-orienting some rooted network that is contained in one of the well-known classes of tree-child, orchard, tree-based or forest-based networks. These results address structural aspects of (multi-)semi-directed networks and pave the way to improved theoretical and computational analyses of such networks, for example, within the development of algebraic evolutionary models that are based on such networks.

q-bio.PE↗

Subtree Distances, Tight Spans and Diversities

Metric embeddings are central to metric theory and its applications. Here we consider embeddings of a different sort: maps from a set to subsets of a metric space so that distances between points are approximated by minimal distances between subsets. Our main result is a characterization of when a set of distances $d(x,y)$ between elements in a set $X$ have a subtree representation, a real tree $T$ and a collection $\{S_x\}_{x \in X}$ of subtrees of~$T$ such that $d(x,y)$ equals the length of the shortest path in~$T$ from a point in $S_x$ to a point in $S_y$ for all $x,y \in X$. The characterization was first established for {\em finite} $X$ by Hirai (2006) using a tight span construction defined for distance spaces, metric spaces without the triangle inequality. To extend Hirai's result beyond finite $X$ we establish fundamental results of tight span theory for general distance spaces, including the surprising observation that the tight span of a distance space is hyperconvex. We apply the results to obtain the first characterization of when a diversity -- a generalization of a metric space which assigns values to all finite subsets of $X$, not just to pairs -- has a tight span which is tree-like.

math.MG↗

When are quarnets sufficient to reconstruct semi-directed phylogenetic networks?

Phylogenetic networks are graphs that are used to represent evolutionary relationships between different taxa. They generalize phylogenetic trees since for example, unlike trees, they permit lineages to combine. Recently, there has been rising interest in semi-directed phylogenetic networks, which are mixed graphs in which certain lineage combination events are represented by directed edges coming together, whereas the remaining edges are left undirected. One reason to consider such networks is that it can be difficult to root a network using real data. In this paper, we consider the problem of when a semi-directed phylogenetic network is defined or encoded by the smaller networks that it induces on the 4-leaf subsets of its leaf set. These smaller networks are called quarnets. We prove that semi-directed binary level-2 phylogenetic networks are encoded by their quarnets, but that this is not the case for level-3. In addition, we prove that the so-called blob tree of a semi-directed binary network, a tree that give the coarse-grained structure of the network, is always encoded by the quarnets of the network. These results are relevant for proving the statistical consistency of programs that are currently being developed for reconstructing phylogenetic networks from practical data, such as the recently developed Squirrel software tool.

q-bio.PE↗

Cherry picking in forests: A new characterization for the unrooted hybrid number of two phylogenetic trees

Phylogenetic networks are a special type of graph which generalize phylogenetic trees and that are used to model non-treelike evolutionary processes such as recombination and hybridization. In this paper, we consider {\em unrooted} phylogenetic networks, i.e. simple, connected graphs $\mathcal{N}=(V,E)$ with leaf set $X$, for $X$ some set of species, in which every internal vertex in $\mathcal{N}$ has degree three. One approach used to construct such phylogenetic networks is to take as input a collection $\mathcal{P}$ of phylogenetic trees and to look for a network $\mathcal{N}$ that contains each tree in $\mathcal{P}$ and that minimizes the quantity $r(\mathcal{N}) = |E|-(|V|-1)$ over all such networks. Such a network always exists, and the quantity $r(\mathcal{N})$ for an optimal network $\mathcal{N}$ is called the hybrid number of $\mathcal{P}$. In this paper, we give a new characterization for the hybrid number in case $\mathcal{P}$ consists of two trees. This characterization is given in terms of a cherry picking sequence for the two trees, although to prove that our characterization holds we need to define the sequence more generally for two forests. Cherry picking sequences have been intensively studied for collections of rooted phylogenetic trees, but our new sequences are the first variant of this concept that can be applied in the unrooted setting. Since the hybrid number of two trees is equal to the well-known tree bisection and reconnection distance between the two trees, our new characterization also provides an alternative way to understand this important tree distance.

math.CO↗

Arboreal networks and their underlying trees

Horizontal gene transfer (HGT) is an important process in bacterial evolution. Current phylogeny-based approaches to capture it cannot however appropriately account for the fact that HGT can occur between bacteria living in different ecological niches. Due to the fact that arboreal networks are a type of multiple-rooted phylogenetic network that can be thought of as a forest of rooted phylogenetic trees along with a set of additional arcs each joining two different trees in the forest, understanding the combinatorial structure of such networks might therefore pave the way to extending current phylogeny-based HGT-inference methods in this direction. A central question in this context is, how can we construct an arboreal network? Answering this question is strongly informed by finding ways to \textit{encode} an arboreal network, that is, breaking up the network into simpler combinatorial structures that, in a well defined sense uniquely determine the network. In the form of triplets, trinets and quarnets such encodings are known for certain types of single-rooted phylogenetic networks. By studying the underlying tree of an arboreal network, we compliment them here with an answer for arboreal networks.

q-bio.PE↗

Phylogenetic trees defined by at most three characters

In evolutionary biology, phylogenetic trees are commonly inferred from a set of characters (partitions) of a collection of biological entities (e.g., species or individuals in a population). Such characters naturally arise from molecular sequences or morphological data. Interestingly, it has been known for some time that any binary phylogenetic tree can be (convexly) defined by a set of at most four characters, and that there are binary phylogenetic trees for which three characters are not enough. Thus, it is of interest to characterise those phylogenetic trees that are defined by a set of at most three characters. In this paper, we provide such a characterisation, in particular proving that a binary phylogenetic tree $T$ is defined by a set of at most three characters precisely if $T$ has no internal subtree isomorphic to a certain tree.

q-bio.PE↗

Orienting undirected phylogenetic networks

This paper studies the relationship between undirected (unrooted) and directed (rooted) phylogenetic networks. We describe a polynomial-time algorithm for deciding whether an undirected nonbinary phylogenetic network, given the locations of the root and reticulation vertices, can be oriented as a directed nonbinary phylogenetic network. Moreover, we characterize when this is possible and show that, in such instances, the resulting directed nonbinary phylogenetic network is unique. In addition, without being given the location of the root and the reticulation vertices, we describe an algorithm for deciding whether an undirected binary phylogenetic network $N$ can be oriented as a directed binary phylogenetic network of a certain class. The algorithm is fixed-parameter tractable (FPT) when the parameter is the level of $N$ and is applicable to classes of directed phylogenetic networks that satisfy certain conditions. As an example, we show that the well-studied class of binary tree-child networks satisfies these conditions.

cs.DS↗

Is this network proper forest-based?

In evolutionary biology, networks are becoming increasingly used to represent evolutionary histories for species that have undergone non-treelike or reticulate evolution. Such networks are essentially directed acyclic graphs with a leaf set that corresponds to a collection of species, and in which non-leaf vertices with indegree 1 correspond to speciation events and vertices with indegree greater than 1 correspond to reticulate events such as gene transfer. Recently forest-based networks have been introduced, which are essentially (multi-rooted) networks that can be formed by adding some arcs to a collection of phylogenetic trees (or phylogenetic forest), where each arc is added in such a way that its ends always lie in two different trees in the forest. In this paper, we consider the complexity of deciding whether or not a given network is proper forest-based, that is, whether it can be formed by adding arcs to some underlying phylogenetic forest which contains the same number of trees as there are roots in the network. More specifically, we show that it can be decided in polynomial time whether or not a binary, tree-child network with $m \ge 2$ roots is proper forest-based in case $m=2$, but that this problem is NP-complete for $m\ge 3$. We also give a fixed parameter tractable (FPT) algorithm for deciding whether or not a network in which every vertex has indegree at most 2 is proper forest-based. A key element in proving our results is a new characterization for when a network with $m$ roots is proper forest-based which is given in terms of the existence of certain $m$-colorings of the vertices of the network.

q-bio.PE↗

Shared ancestry graphs and symbolic arboreal maps

A network $N$ on a finite set $X$, $|X|\geq 2$, is a connected directed acyclic graph with leaf set $X$ in which every root in $N$ has outdegree at least 2 and no vertex in $N$ has indegree and outdegree equal to 1; $N$ is arboreal if the underlying unrooted, undirected graph of $N$ is a tree. Networks are of interest in evolutionary biology since they are used, for example, to represent the evolutionary history of a set $X$ of species whose ancestors have exchanged genes in the past. For $M$ some arbitrary set of symbols, $d:{X \choose 2} \to M \cup \{\odot\}$ is a symbolic arboreal map if there exists some arboreal network $N$ whose vertices with outdegree two or more are labelled by elements in $M$ and so that $d(\{x,y\})$, $\{x,y\} \in {X \choose 2}$, is equal to the label of the least common ancestor of $x$ and $y$ in $N$ if this exists and $\odot$ else. Important examples of symbolic arboreal maps include the symbolic ultrametrics, which arise in areas such as game theory, phylogenetics and cograph theory. In this paper we show that a map $d:{X \choose 2} \to M \cup \{\odot\}$ is a symbolic arboreal map if and only if $d$ satisfies certain 3- and 4-point conditions and the graph with vertex set $X$ and edge set consisting of those pairs $\{x,y\} \in {X \choose 2}$ with $d(\{x,y\}) \neq \odot$ is Ptolemaic. To do this, we introduce and prove a key theorem concerning the shared ancestry graph for a network $N$ on $X$, where this is the graph with vertex set $X$ and edge set consisting of those $\{x,y\} \in {X \choose 2}$ such that $x$ and $y$ share a common ancestor in $N$. In particular, we show that for any connected graph $G$ with vertex set $X$ and edge clique cover $K$ in which there are no two distinct sets in $K$ with one a subset of the other, there is some network with $|K|$ roots and leaf set $X$ whose shared ancestry graph is $G$.

math.CO↗

Autopolyploidy, allopolyploidy, and phylogenetic networks with horizontal arcs

Polyploidization is an evolutionary process by which a species acquires multiple copies of its complete set of chromosomes. The reticulate nature of the signal left behind by it means that phylogenetic networks offer themselves as a framework to reconstruct the evolutionary past of species affected by it. The main strategy for doing this is to first construct a so called multiple-labelled tree and to then somehow derive such a network from it. The following question therefore arises: How much can be said about that past if such a tree is not readily available? By viewing a polyploid dataset as a certain vector which we call a ploidy (level) profile we show that, among other results, there always exists a phylogenetic network in the form of a beaded phylogenetic tree with additional arcs that realizes a given ploidy profile. Intriguingly, the two end vertices of almost all of these additional arcs can be interpreted as having co-existed in time thereby adding biological realism to our network, a feature that is, in general, not enjoyed by phylogenetic networks. In addition, we show that our network may be viewed as a generator of ploidy profile space, a novel concept similar to phylogenetic tree space that we introduce to be able to compare phylogenetic networks that realize one and the same ploidy profile. We illustrate our findings in terms of a publicly available Viola dataset.

q-bio.PE↗

Diversities and the Generalized Circumradius

The generalized circumradius of a set of points $A \subseteq \mathbb{R}^d$ with respect to a convex body $K$ equals the minimum value of $λ\geq 0$ such that $A$ is contained in a translate of $λK$. Each choice of $K$ gives a different function on the set of bounded subsets of $\mathbb{R}^d$; we characterize which functions can arise in this way. Our characterization draws on the theory of diversities, a recently introduced generalization of metrics from functions on pairs to functions on finite subsets. We additionally investigate functions which arise by restricting the generalised circumradius to a finite subset of $\mathbb{R}^d$. We obtain elegant characterizations in the case that $K$ is a simplex or parallelotope.

math.MG↗

SPRINT: A fast, new software tool for reconstructing the evolutionary past of polyploid datasets

Polyploidization is an important evolutionary process which affects organisms ranging from plants to fish and fungi. The signal left behind by it is in the form of a species' ploidy level (number of complete chromosome sets found in a cell) which is inherently non-treelike. Currently available tools for reconstructing the evolutionary past of a polyploid dataset generally start with a multi-labelled tree obtained for a dataset of interest and then derive a (phylogenetic) network from that tree in some way that reflects that past by interpreting the networks's vertices of indegree at least two as polyploidization events. Since obtaining such a tree can be computationally expensive it is paramount to have alternative approaches available that allow one to shed light into the reticulate evolutionary past of a polyploid dataset. SPRINT aims to reconstruct the evolutionary past of a polyploid dataset in terms of a binary network which realises the dataset's ploidy profile (vector of ploidy levels of the dataset's taxa) and requires the fewest number of polyploidization events. It does this by representing the ploidy level of a species x in terms of the number of directed paths from the root of the network to the leaf of the network labelled by x. SPRINT is distributed on GitHub: https://github.com/lmaher1/SPRINT.

q-bio.PE↗

The hybrid number of a ploidy profile

Polyploidization, whereby an organism inherits multiple copies of the genome of their parents, is an important evolutionary event that has been observed in plants and animals. One way to study such events is in terms of the ploidy number of the species that make up a dataset of interest. It is therefore natural to ask: How much information about the evolutionary past of the set of species that form a dataset can be gleaned from the ploidy numbers of the species? To help answer this question, we introduce and study the novel concept of a ploidy profile which allows us to formalize it in terms of a multiplicity vector indexed by the species the dataset is comprised of. Using the framework of a phylogenetic network, we present a closed formula for computing the hybrid number (i.e. the minimal number of polyploidization events required to explain a ploidy profile) of a large class of ploidy profiles. This formula relies on the construction of a certain phylogenetic network from the simplification sequence of a ploidy profile and the hybrid number of the ploidy profile with which this construction is initialized. Both of them can be computed easily in case the ploidy numbers that make up the ploidy profile are not too large. To help illustrate the applicability of our approach, we apply it to a simplified version of a publicly available Viola dataset.

q-bio.PE↗

Forest-based networks

In evolutionary studies it is common to use phylogenetic trees to represent the evolutionary history of a set of species. However, in case the transfer of genes or other genetic information between the species or their ancestors has occurred in the past, a tree may not provide a complete picture of their history. In such cases,tree-based phylogenetic networks can provide a useful, more refined representation of the species evolution. Such a network is essentially a phylogenetic tree with some arcs added between the tree edges so as to represent reticulate events such as gene transfer. Even so, this model does not permit the representation of evolutionary scenarios where reticulate events have taken place between different subfamilies or lineages of species. To represent such scenarios, in this paper we introduce the notion of a forest-based phylogenetic network, that is, a collection of leaf-disjoint phylogenetic trees on a set of species with arcs added between the edges of distinct trees within the collection. Forest-based networks include the recently introduced class of overlaid species forests which are used to model introgression. As we shall see, even though the definition of forest-based networks is closely related to that of tree-based networks, they lead to new mathematical theory which complements that of tree-based networks. As well as studying the relationship of forest-based networks with other classes of phylogenetic networks, such as tree-child networks and universal tree-based networks, we present some characterizations of some special classes of forest-based networks. We expect that our results will be useful for developing new models and algorithms to understand reticulate evolution, such as gene transfer between collections of bacteria that live in different environments.

math.CO↗

The space of equidistant phylogenetic cactuses

We introduce and investigate the space of \emph{equidistant} $X$-\emph{cactuses}. These are rooted, arc weighted, phylogenetic networks with leaf set $X$, where $X$ is a finite set of species, and all leaves have the same distance from the root. The space contains as a subset the space of ultrametric trees on $X$ that was introduced by Gavryushkin and Drummond. We show that equidistant-cactus space is a CAT(0)-metric space which implies, for example, that there are unique geodesic paths between points. As a key step to proving this, we present a combinatorial result concerning \emph{ranked} rooted $X$-cactuses. In particular, we show that such networks can be encoded in terms of a pairwise compatibility condition arising from a poset of collections of pairs of subsets of $X$ that satisfy certain set-theoretic properties. As a corollary, we also obtain an encoding of ranked, rooted $X$-trees in terms of partitions of $X$, which provides an alternative proof that the space of ultrametric trees on $X$ is CAT(0). As with spaces of phylogenetic trees, we expect that our results should provide the basis for and new directions in performing statistical analyses for collections of phylogenetic networks with arc lengths.

q-bio.PE↗