SearcharxivSearch

arXiv subjects

Sagi Snir

Publications and source records attributed to Sagi Snir.

8 recordsLinked to original sources

Simultaneous separation in bounded degree trees

It follows from a classical result of Jordan that every tree with maximum degree at most $r$ containing a vertex set labeled by $[n]$, has a single-edge cut which separates two subsets $A,B \subset [n]$ for which $\min\{|A|,|B|\} \ge (n-1)/r$. Motivated by the tree dissimilarity problem in phylogenetics, we consider the case of separating vertex sets of {\em several} trees: Given $k$ trees with maximum degree at most $r$, containing a common vertex set labeled by $[n]$, we ask for a single-edge cut in each tree which maximizes $min\{|A|,|B|\}$ where $A,B \subset [n]$ are separated by the corresponding cut at each tree. Denoting this maximum by $f(r,k,n)$ and considering the limit $f(r,k) = \lim_{n \rightarrow \infty} f(r,k,n)/n$ (which is shown to always exist) we determine that $f(r,2)=\frac{1}{2r}$ and determine that $f(3,3)=\frac{2}{27}$, which is already quite intricate. The case $r=3$ is especially interesting in phylogenetics and our result implies that any two (three) binary phylogenetic trees over $n$ taxa have a split at each tree which separates two taxa sets of order at least $n/6$ (resp. $2n/27$), and these bounds are asymptotically tight.

math.CO

On the quartet distance given partial information

Let $T$ be an arbitrary phylogenetic tree with $n$ leaves. It is well-known that the average quartet distance between two assignments of taxa to the leaves of $T$ is $\frac 23 \binom{n}{4}$. However, a longstanding conjecture of Bandelt and Dress asserts that $(\frac 23 +o(1))\binom{n}{4}$ is also the {\em maximum} quartet distance between two assignments. While Alon, Naves, and Sudakov have shown this indeed holds for caterpillar trees, the general case of the conjecture is still unresolved. A natural extension is when partial information is given: the two assignments are known to coincide on a given subset of taxa. The partial information setting is biologically relevant as the location of some taxa (species) in the phylogenetic tree may be known, and for other taxa it might not be known. What can we then say about the average and maximum quartet distance in this more general setting? Surprisingly, even determining the {\em average} quartet distance becomes a nontrivial task in the partial information setting and determining the maximum quartet distance is even more challenging, as these turn out to be dependent of the structure of $T$. In this paper we prove nontrivial asymptotic bounds that are sometimes tight for the average quartet distance in the partial information setting. We also show that the Bandelt and Dress conjecture does not generally hold under the partial information setting. Specifically, we prove that there are cases where the average and maximum quartet distance substantially differ.

q-bio.PE

Multi-SpaM: a Maximum-Likelihood approach to Phylogeny reconstruction based on Multiple Spaced-Word Matches

Motivation: Word-based or `alignment-free' methods for phylogeny reconstruction are much faster than traditional approaches, but they are generally less accurate. Most of these methods calculate pairwise distances for a set of input sequences, for example from word frequencies, from so-called spaced-word matches or from the average length of common substrings. Results: In this paper, we propose the first word-based approach to tree reconstruction that is based on multiple sequence comparison and Maximum Likelihood. Our algorithm first samples small, gap-free alignments involving four taxa each. For each of these alignments, it then calculates a quartet tree and, finally, the program Quartet MaxCut is used to infer a super tree topology for the full set of input taxa from the calculated quartet trees. Experimental results show that trees calculated with our approach are of high quality. Availability: The source code of the program is available at https://github.com/tdencker/multi-SpaM Contact: thomas.dencker@stud.uni-goettingen.de

q-bio.PE

Recovering the tree-like trend of evolution despite extensive lateral genetic transfer: A probabilistic analysis

Lateral gene transfer (LGT) is a common mechanism of non-vertical evolution where genetic material is transferred between two more or less distantly related organisms. It is particularly common in bacteria where it contributes to adaptive evolution with important medical implications. In evolutionary studies, LGT has been shown to create widespread discordance between gene trees as genomes become mosaics of gene histories. In particular, the Tree of Life has been questioned as an appropriate representation of bacterial evolutionary history. Nevertheless a common hypothesis is that prokaryotic evolution is primarily tree-like, but that the underlying trend is obscured by LGT. Extensive empirical work has sought to extract a common tree-like signal from conflicting gene trees. Here we give a probabilistic perspective on the problem of recovering the tree-like trend despite LGT. Under a model of randomly distributed LGT, we show that the species phylogeny can be reconstructed even in the presence of surprisingly many (almost linear number of) LGT events per gene tree. Our results, which are optimal up to logarithmic factors, are based on the analysis of a robust, computationally efficient reconstruction method and provides insight into the design of such methods. Finally we show that our results have implications for the discovery of highways of gene sharing.

math.PR

Universal pacemaker of genome evolution

Molecular clock (MC) is a central concept of molecular evolution according to which each gene evolves at a characteristic, near constant rate. Numerous evolutionary studies have demonstrated the validity of MC but also have shown that MC is substantially overdispersed, i.e. lineage-specific deviations of the evolutionary rate of the given gene from the clock greatly exceed the expectation from the sampling error. A fundamental observation of comparative genomics that appears to complement the MC is that the distribution of evolution rates across orthologous genes in pairs of related genomes remains virtually unchanged throughout the evolution of life, from bacteria to mammals. The conservation of this distribution implies that the relative evolution rates of all genes remain nearly constant, or in other words, that evolutionary rates of different genes are strongly correlated within each evolving genome. We hypothesized that this correlation is not a simple consequence of MC but could be better explained by a model we dubbed Universal PaceMaker (UPM) of genome evolution. The UPM model posits that the rate of evolution changes synchronously across genome-wide sets of genes in all evolving lineages. We sought to differentiate between the MC and UPM models by fitting thousands of phylogenetic trees for bacterial and archaeal genes to supertrees that reflect the dominant trend of vertical descent in the evolution of archaea and bacteria and that were constrained according to the two models. The goodness of fit for the UPM model was better than the fit for the MC model, with overwhelming statistical significance. These results reveal a universal pacemaker of genome evolution that could have been in operation throughout the history of life.

q-bio.PE

Hadamard Conjugation for the Kimura 3ST Model: Combinatorial Proof using Pathsets

In most stochastic models of molecular sequence evolution the probability of each possible pattern of homologous characters at a site is estimated numerically. However in the case of Kimura's three-substitution-types (K3ST) model, these probabilities can be expressed analytically by Hadamard conjugation as a function of the phylogeny T and the substitution probabilities on each edge of T, together with an analytic inverse function. In this paper we produce a direct proof of these results, using pathset distances which generalise pairwise distances between sequences. This interpretation allows us to apply Hadamard conjugation to a number of topical problems in the mathematical analysis of sequence evolution.

q-bio.PE

Efficient Approximation of Convex Recolorings

A coloring of a tree is convex if the vertices that pertain to any color induce a connected subtree; a partial coloring (which assigns colors to some of the vertices) is convex if it can be completed to a convex (total) coloring. Convex coloring of trees arise in areas such as phylogenetics, linguistics, etc. eg, a perfect phylogenetic tree is one in which the states of each character induce a convex coloring of the tree. Research on perfect phylogeny is usually focused on finding a tree so that few predetermined partial colorings of its vertices are convex. When a coloring of a tree is not convex, it is desirable to know "how far" it is from a convex one. In [19], a natural measure for this distance, called the recoloring distance was defined: the minimal number of color changes at the vertices needed to make the coloring convex. This can be viewed as minimizing the number of "exceptional vertices" w.r.t. to a closest convex coloring. The problem was proved to be NP-hard even for colored string. In this paper we continue the work of [19], and present a 2-approximation algorithm of convex recoloring of strings whose running time O(cn), where c is the number of colors and n is the size of the input, and an O(cn^2)-time 3-approximation algorithm for convex recoloring of trees.

cs.DS

Maximum Likelihood Jukes-Cantor Triplets: Analytic Solutions

Complex systems of polynomial equations have to be set up and solved algebraically in order to obtain analytic solutions for maximum likelihood on phylogenetic trees. This has restricted the types of systems previously resolved to the simplest models - three and four taxa under a molecular clock, with just two state characters. In this work we give, for the first time, analytic solutions for a family of trees with four state characters, like normal DNA or RNA. The model of substitution we use is the Jukes-Cantor model, and the trees are on three taxa under molecular clock, namely rooted triplets. We employ a number of approaches and tools to solve this system: Spectral methods (Hadamard conjugation), a new representation of variables (the path-set spectrum), and algebraic geometry tools (the resultant of two polynomials). All these, combined with heavy application of computer algebra packages (Maple), let us derive the desired solution.

q-bio.PE