SearcharxivSearch

arXiv subjects

Filippo Disanto

Publications and source records attributed to Filippo Disanto.

At least 19 recordsLinked to original sources

The distributions under two species-tree models of the total number of ancestral configurations for matching gene trees and species trees

Given a gene-tree labeled topology $G$ and a species tree $S$, the "ancestral configurations" at an internal node $k$ of $S$ represent the combinatorially different sets of gene lineages that can be present at $k$ when all possible realizations of $G$ in $S$ are considered. Ancestral configurations have been introduced as a data structure for evaluating the conditional probability of a gene-tree labeled topology given a species tree, and their enumeration assists in describing the complexity of this computation. In the case that the gene-tree labeled topology $G=t$ matches that of the species tree $S$, by techniques of analytic combinatorics, we study distributional properties of the "total" number of ancestral configurations measured across the different nodes of a random labeled topology $t$ selected under the uniform and the Yule probability models. Under both of these probabilistic scenarios, we show that the total number $T_n$ of ancestral configurations of a random labeled topology of $n$ taxa asymptotically follows a lognormal distribution. Over uniformly distributed labeled topologies, the asymptotic growth of the mean and the variance of $T_n$ are found to satisfy $\mathbb{E}_{\rm U}[T_n] \sim 2.449 \cdot 1.333^n$ and $\mathbb{V}_{\rm U}[T_n] \sim 5.050 \cdot 1.822^n$, respectively. Under the Yule model, which assigns higher probabilities to more balanced labeled topologies, we obtain the mean $\mathbb{E}_{\rm Y}[T_n] \sim 1.425^n$ and the variance $\mathbb{V}_{\rm Y}[T_n] \sim 2.045^n$.

math.PR

Distribution of external branch lengths in Yule trees

The Yule branching process is a classical model for the random generation of gene tree topologies in population genetics. It generates binary ranked trees -- also called "histories" -- with a finite number $n$ of leaves. We study the lengths $\ell_1 > \ell_2 > ... > \ell_k > ...$ of the external branches of a Yule generated random history of size $n$, where the length of an external branch is defined as the rank of its parent node. When $n \rightarrow \infty$, we show that the random variable $\ell_k$, once rescaled as $\frac{n-\ell_k}{\sqrt{n/2}}$, follows a $\chi$-distribution with $2k$ degrees of freedom, with mean $\mathbb E(\ell_k) \sim n$ and variance $\mathbb V(\ell_k) \sim n \big(k-\frac{\pi k^2}{16^k} \binom{2k}{k}^2\big)$. Our results contribute to the study of the combinatorial features of Yule generated gene trees, in which external branches are associated with singleton mutations affecting individual gene copies.

math.PR

Enumeration of binary trees compatible with a perfect phylogeny

Evolutionary models used for describing molecular sequence variation suppose that at a non-recombining genomic segment, sequences share ancestry that can be represented as a genealogy--a rooted, binary, timed tree, with tips corresponding to individual sequences. Under the infinitely-many-sites mutation model, mutations are randomly superimposed along the branches of the genealogy, so that every mutation occurs at a chromosomal site that has not previously mutated; if a mutation occurs at an interior branch, then all individuals descending from that branch carry the mutation. The implication is that observed patterns of molecular variation from this model impose combinatorial constraints on the hidden state space of genealogies. In particular, observed molecular variation can be represented in the form of a perfect phylogeny, a tree structure that fully encodes the mutational differences among sequences. For a sample of n sequences, a perfect phylogeny might not possess n distinct leaves, and hence might be compatible with many possible binary tree structures that could describe the evolutionary relationships among the n sequences. Here, we investigate enumerative properties of the set of binary ranked and unranked tree shapes that are compatible with a perfect phylogeny, and hence, the binary ranked and unranked tree shapes conditioned on an observed pattern of mutations under the infinitely-many-sites mutation model. We provide a recursive enumeration of these shapes. We consider both perfect phylogenies that can be represented as binary and those that are multifurcating. The results have implications for computational aspects of the statistical inference of evolutionary parameters that underlie sets of molecular sequences.

q-bio.PE

The distributions under two species-tree models of the number of root ancestral configurations for matching gene trees and species trees

For a pair consisting of a gene tree and a species tree, the ancestral configurations at an internal node of the species tree are the distinct sets of gene lineages that can be present at that node. Ancestral configurations appear in computations of gene tree probabilities under evolutionary models conditional on fixed species trees, and the enumeration of root ancestral configurations -- ancestral configurations at the root of the species tree -- assists in describing the complexity of these computations. In the case that the gene tree matches the species tree in topology, we study the distribution of the number of root ancestral configurations of a random labeled tree topology under each of two models.

math.CO

Enumeration of ancestral configurations for matching gene trees and species trees

Given a gene tree and a species tree, ancestral configurations represent the combinatorially distinct sets of gene lineages that can reach a given node of the species tree. They have been introduced as a data structure for use in the recursive computation of the conditional probability under the multispecies coalescent model of a gene tree topology given a species tree, the cost of this computation being affected by the number of ancestral configurations of the gene tree in the species tree. For matching gene trees and species trees, we obtain enumerative results on ancestral configurations. We study ancestral configurations in balanced and unbalanced families of trees determined by a given seed tree, showing that for seed trees with more than one taxon, the number of ancestral configurations increases for both families exponentially in the number of taxa $n$. For fixed $n$, the maximal number of ancestral configurations tabulated at the species tree root node and the largest number of labeled histories possible for a labeled topology occur for trees with precisely the same unlabeled shape. For ancestral configurations at the root, the maximum increases with $k_0^n$, where $k_0 \approx 1.5028$ is a quadratic recurrence constant. Under a uniform distribution over the set of labeled trees of given size, the mean number of root ancestral configurations grows with $\sqrt{3/2}(4/3)^n$ and the variance with approximately $1.4048(1.8215)^n$. The results provide a contribution to the combinatorial study of gene trees and species trees.

q-bio.PE

Asymptotic properties of the number of matching coalescent histories for caterpillar-like families of species trees

Coalescent histories provide lists of species tree branches on which gene tree coalescences can take place, and their enumerative properties assist in understanding the computational complexity of calculations central in the study of gene trees and species trees. Here, we solve an enumerative problem left open by Rosenberg (IEEE/ACM Transactions on Computational Biology and Bioinformatics 10: 1253-1262, 2013) concerning the number of coalescent histories for gene trees and species trees with a matching labeled topology that belongs to a generic caterpillar-like family. By bringing a generating function approach to the study of coalescent histories, we prove that for any caterpillar-like family with seed tree $t$, the sequence $(h_n)_{n\geq 0}$ describing the number of matching coalescent histories of the $n$th tree of the family grows asymptotically as a constant multiple of the Catalan numbers. Thus, $h_n \sim \beta_t c_n$, where the asymptotic constant $\beta_t > 0$ depends on the shape of the seed tree $t$. The result extends a claim demonstrated only for seed trees with at most 8 taxa to arbitrary seed trees, expanding the set of cases for which detailed enumerative properties of coalescent histories can be determined. We introduce a procedure that computes from $t$ the constant $\beta_t$ as well as the algebraic expression for the generating function of the sequence $(h_n)_{n\geq 0}$.

q-bio.PE

Coalescent histories for lodgepole species trees

Coalescent histories are combinatorial structures that describe for a given gene tree and species tree the possible lists of branches of the species tree on which the gene tree coalescences take place. Properties of the number of coalescent histories for gene trees and species trees affect a variety of probabilistic calculations in mathematical phylogenetics. Exact and asymptotic evaluations of the number of coalescent histories, however, are known only in a limited number of cases. Here we introduce a particular family of species trees, the \emph{lodgepole} species trees $(\lambda_n)_{n\geq 0}$, in which tree $\lambda_n$ has $m=2n+1$ taxa. We determine the number of coalescent histories for the lodgepole species trees, in the case that the gene tree matches the species tree, showing that this number grows with $m!!$ in the number of taxa $m$. This computation demonstrates the existence of tree families in which the growth in the number of coalescent histories is faster than exponential. Further, it provides a substantial improvement on the lower bound for the ratio of the largest number of matching coalescent histories to the smallest number of matching coalescent histories for trees with $m$ taxa, increasing a previous bound of $(\sqrt{\pi} / 32)[(5m-12)/(4m-6)] m \sqrt{m}$ to $[ \sqrt{m-1}/(4 \sqrt{e}) ]^{m}$. We discuss the implications of our enumerative results for phylogenetic computations.

q-bio.PE

On the number of ranked species trees producing anomalous ranked gene trees

Analysis of probability distributions conditional on species trees has demonstrated the existence of anomalous ranked gene trees (ARGTs), ranked gene trees that are more probable than the ranked gene tree that accords with the ranked species tree. Here, to improve the characterization of ARGTs, we study enumerative and probabilistic properties of two classes of ranked labeled species trees, focusing on the presence or avoidance of certain subtree patterns associated with the production of ARGTs. We provide exact enumerations and asymptotic estimates for cardinalities of these sets of trees, showing that as the number of species increases without bound, the fraction of all ranked labeled species trees that are ARGT-producing approaches 1. This result extends beyond earlier existence results to provide a probabilistic claim about the frequency of ARGTs.

q-bio.PE

Yule-generated trees constrained by node imbalance

The Yule process generates a class of binary trees which is fundamental to population genetic models and other applications in evolutionary biology. In this paper, we introduce a family of sub-classes of ranked trees, called Omega-trees, which are characterized by imbalance of internal nodes. The degree of imbalance is defined by an integer 0 <= w. For caterpillars, the extreme case of unbalanced trees, w = 0. Under models of neutral evolution, for instance the Yule model, trees with small w are unlikely to occur by chance. Indeed, imbalance can be a signature of permanent selection pressure, such as observable in the genealogies of certain pathogens. From a mathematical point of view it is interesting to observe that the space of Omega-trees maintains several statistical invariants although it is drastically reduced in size compared to the space of unconstrained Yule trees. Using generating functions, we study here some basic combinatorial properties of Omega-trees. We focus on the distribution of the number of subtrees with two leaves. We show that expectation and variance of this distribution match those for unconstrained trees already for very small values of w.

q-bio.PE

On the maximal weight of $(p,q)$-ary chain partitions with bounded parts

A $(p,q)$-ary chain is a special type of chain partition of integers with parts of the form $p^aq^b$ for some fixed integers $p$ and $q$. In this note, we are interested in the maximal weight of such partitions when their parts are distinct and cannot exceed a given bound $m$. Characterizing the cases where the greedy choice fails, we prove that this maximal weight is, as a function of $m$, asymptotically independent of $\max(p,q)$, and we provide an efficient algorithm to compute it.

math.NT

On the sub-permutations of pattern avoiding permutations

There is a deep connection between permutations and trees. Certain sub-structures of permutations, called sub-permutations, bijectively map to sub-trees of binary increasing trees. This opens a powerful tool set to study enumerative and probabilistic properties of sub-permutations and to investigate the relationships between 'local' and 'global' features using the concept of pattern avoidance. First, given a pattern {\mu}, we study how the avoidance of {\mu} in a permutation {\pi} affects the presence of other patterns in the sub-permutations of {\pi}. More precisely, considering patterns of length 3, we solve instances of the following problem: given a class of permutations K and a pattern {\mu}, we ask for the number of permutations $\pi \in Av_n(\mu)$ whose sub-permutations in K satisfy certain additional constraints on their size. Second, we study the probability for a generic pattern to be contained in a random permutation {\pi} of size n without being present in the sub-permutations of {\pi} generated by the entry $1 \leq k \leq n$. These theoretical results can be useful to define efficient randomized pattern-search procedures based on classical algorithms of pattern-recognition, while the general problem of pattern-search is NP-complete.

math.CO

A partial order structure on interval orders

We introduce a partial order structure on the set of interval orders of a given size, and prove that such a structure is in fact a lattice. We also provide a way to compute meet and join inside this lattice. Finally, we show that, if we restrict to series parallel interval order, what we obtain is the classical Tamari poset.

math.CO

Unbalanced subtrees in binary rooted ordered and un-ordered trees

Binary rooted trees, both in the ordered and in the un-ordered case, are well studied structures in the field of combinatorics. The aim of this work is to study particular patterns in these classes of trees. We consider completely unbalanced subtrees, where unbalancing is measured according to the so-called Colless's index. The size of the biggest unbalanced subtree becomes then a new parameter with respect to which we find several enumerations.

math.CO

Andre' permutations, right-to-left and left-to-right minima

We provide enumerative results concerning right-to-left minima and left- to-right minima in Andre' permutations of the first and second kind. For both the two kinds, the distribution of right-to-left and left-to-right minima is the same. We provide generating functions and associated asymptotics results. Our approach is based on the tree-structure of Andre' permutations.

math.CO

Exact enumeration of cherries and pitchforks in ranked trees under the coalescent model

We consider exact enumerations and probabilistic properties of ranked trees when generated under the random coalescent process. Using a new approach, based on generating functions, we derive several statistics such as the exact probability of finding k cherries in a ranked tree of fixed size n. We then extend our method to consider also the number of pitchforks. We find a recursive formula to calculate the joint and conditional probabilities of cherries and pitch- forks when the size of the tree is fixed.

math.CO

Catalan structures and Catalan pairs

A Catalan pair is a pair of binary relations (S,R) satisfying certain axioms. These objects are enumerated by the well-known Catalan numbers, and have been introduced with the aim of giving a common language to most of the structures counted by Catalan numbers. Here, we give a simple method to pass from the recursive definition of a generic Catalan structure to the recursive definition of the Catalan pair on the same structure, thus giving an automatic way to interpret Catalan structures in terms of Catalan pairs. We apply our method to many well-known Catalan structures, focusing on the meaning of the relations S and R in each considered case.

cs.DM

Catalan lattices on series parallel interval orders

Using the notion of series parallel interval order, we propose a unified setting to describe Dyck lattices and Tamari lattices (two well known lattice structures on Catalan objects) in terms of basic notions of the theory of posets. As a consequence of our approach, we find an extremely simple proof of the fact that the Dyck order is a refinement of the Tamari one. Moreover, we provide a description of both the weak and the strong Bruhat order on 312-avoiding permutations, by recovering the proof of the fact that they are isomorphic to the Tamari and the Dyck order, respectively; our proof, which simplifies the existing ones, relies on our results on series parallel interval orders.

math.CO

Catalan numbers and relations

We define the notion of a Catalan pair (which is a pair of binary relations (S,R) satisfying certain axioms) with the aim of giving a common language to most of the combinatorial interpretations of Catalan numbers. We show, in particular, that the second component R uniquely determines the pair, and we give a characterization of R in terms of forbidden configurations. We also propose some generalizations of Catalan pairs arising from some slight modifications of (some of the) axioms.

math.CO