SearcharxivSearch

arXiv subjects

Joseph Rusinko

Publications and source records attributed to Joseph Rusinko.

10 recordsLinked to original sources

Statistical learning with phylogenetic network invariants

Phylogenetic networks provide a means of describing the evolutionary history of sets of species believed to have undergone hybridization or gene flow during their evolution. The mutation process for a set of such species can be modeled as a Markov process on a phylogenetic network. Previous work has shown that a site-pattern probability distributions from a Jukes-Cantor phylogenetic network model must satisfy certain algebraic invariants. As a corollary, aspects of the phylogenetic network are theoretically identifiable from site-pattern frequencies. In practice, because of the probabilistic nature of sequence evolution, the phylogenetic network invariants will rarely be satisfied, even for data generated under the model. Thus, using network invariants for inferring phylogenetic networks requires some means of interpreting the residuals, or deviations from zero, when observed site-pattern frequencies are substituted into the invariants. In this work, we propose a method of utilizing invariant residuals and support vector machines to infer 4-leaf level-one phylogenetic networks, from which larger networks can be reconstructed. Given data for a set of species, the support vector machine is first trained on model data to learn the patterns of residuals corresponding to different network structures to classify the network that produced the data. We demonstrate the performance of our method on simulated data from the specified model and primate data.

q-bio.PE

Constructing Semi-Directed Level-1 Phylogenetic Networks from Quarnets

Semi-directed networks provide a graphical structure for describing the evolutionary history of organisms in the presence of hybridization. We introduce two algorithms for reconstructing semi-directed level-1 phylogenetic networks from their complete set of 4-leaf subnetworks, known as quarnets. The sequential algorithm begins with a single quarnet and adds one leaf at a time until the entire network has been reconstructed. The cherry-blob algorithm is a novel approach inspired by cherry-picking techniques on trees

q-bio.GN

Phylogenetic Networks

Phylogenetics is the study of the evolutionary relationships between organisms. One of the main challenges in the field is to take biological data for a group of organisms and to infer an evolutionary tree, a graph that represents these relationships. Developing practical and efficient methods for inferring phylogenetic trees has lead to a number of interesting mathematical questions across a variety of fields. However, due to hybridization and gene flow, a phylogenetic network may be a better representation of the evolutionary history of some groups of organisms. In this chapter, we introduce some of the basic concepts in phylogenetics, and present related undergraduate research projects on phylogenetic networks that touch on areas of graph theory and abstract algebra. In the first section, we describe several open research questions related to the combinatorics of phylogenetic networks. In the second, we describe problems related to understanding phylogenetic statistical models as algebraic varieties.

q-bio.PE

Phylogenetic Derivative: A Tool for Assessing Local Tree Reconstruction in the Presence of Recombination

Recently, much attention has been given to understanding recombination events along a chromosome in a variety of field. For instance, many population genetics problems are limited by the inaccuracy of inferred evolutionary histories of chromosomes sampled randomly from a population. This evolutionary history differs among genomic locations as an artifact of recombination events along a chromosome. Thus, much recent attention has been focused on identifying these recombination points. However, many proposed methods either make simplifying, but unrealistic, assumptions about recombination along a chromosome, or are unable to scale to large genome-wide data like what has become commonplace in statistical genetics. Here, we introduce a \emph{phylogenetic derivative} to describe the relatedness of neighboring trees along a chromosome. This phylogenetic derivative is a computationally efficient, flexible metric that can be also be used assess the prevalence of recombination across a chromosome. These proposed methods are tested and perform well in analyzing both simulated data and a real mouse data set.

q-bio.QM

Efficient Quartet Representations of Trees and Applications to Supertree and Summary Methods

Quartet trees displayed by larger phylogenetic trees have long been used as inputs for species tree and supertree reconstruction. Computational constraints prevent the use of all displayed quartets in many practical problems due to the number of taxa. We introduce the notion of an Efficient Quartet System (EQS) to represent a phylogenetic tree with a subset of the quartets displayed by the tree. We show mathematically that the set of quartets obtained from a tree via an EQS contains all of the combinatorial information of the tree itself. Using performance tests on simulated datasets, we also demonstrate that using an EQS to reduce the number of quartets in pipelines for summary methods of species tree inference and supertree inference results in only small reductions in accuracy.

q-bio.PE

Species tree estimation using Neighbor Joining

Recent theoretical work has demonstrated that Neighbor Joining applied to concatenated DNA sequences is a statistically consistent method of species tree reconstruction. This brief note compares the accuracy of this approach to other popular statistically consistent species tree reconstruction algorithms including ASTRAL-II Neighbor Joining using average gene-tree internode distances (NJst) and SVD-Quartets+PAUP*, as well as concatenation using maximum likelihood (RaxML). We find that the faster Neighbor Joining, applied to concatenated sequences, is among the most effective of these methods for accurate species tree reconstruction.

q-bio.PE

Modeling the distribution of distance data in Euclidean space

Phylogenetic inference-the derivation of a hypothesis for the common evolutionary history of a group of species- is an active area of research at the intersection of biology, computer science, mathematics, and statistics. One assumes the data contains a phylogenetic signal that will be recovered with varying accuracy due to the quality of the method used, and the quality of the data. The input for distance-based inference methods is an element of a Euclidean space with coordinates indexed by the pairs of organisms. For several algorithms there exists a subdivision of this space into polyhedral cones such that inputs in the same cone return the same tree topology. The geometry of these cones has been used to analyze the inference algorithms. In this chapter, we model how input data points drawn from DNA sequences are distributed throughout Euclidean space in relation to the space of tree metrics, which in turn can also be described as a collection of polyhedral cones.

q-bio.PE

H-Representation of the Kimura-3 Polytope

Given a group-based Markov model on a tree, one can compute the vertex representation of a polytope which describes the associated toric variety. The half-space representation, however, is not easily computable. In the case of $\mathbb{Z}_2$ or $\mathbb{Z}_2\times\mathbb{Z}_2$, these polytopes have applications in the field of phylogenetics. We provide a half-space representation for the $m$-claw tree where $G=\mathbb{Z}_2\times\mathbb{Z}_2$, which corresponds to the Kimura-3 model of evolution.

math.AG

Combinatorics of Linked Systems of Quartet Trees

We apply classical quartet techniques to the problem of phylogenetic decisiveness and find a value $k$ such that all collections of at least $k$ quartets are decisive. Moreover, we prove that this bound is optimal and give a lower-bound on the probability that a collection of quartets is decisive.

q-bio.QM

Combinatorics of k-Interval Cospeciation for Cophylogeny

We show that the cophylogenetic distance, k-interval cospeciation, is distinct from other metrics and accounts for global congruence between locally incongruent trees. The growth of the neighborhood of trees which satisfy the largest possible k-interval cospeciation with a given tree indicates that k-interval cospeciation is useful for analyzing simulated data.

math.CO