Searcharxiv⌕ Search

arXiv subjects

Barbara R. Holland

Publications and source records attributed to Barbara R. Holland.

16 recordsLinked to original sources

Quasi-birth-and-death processes evolving within trees: Applications to comparative phylogenetics

We consider a quasi-birth-and-process (QBD) that duplicates itself at some fixed times within a tree that contains information about duplication times and potentially partially observed states. We analyse a continuous trait by discretising it to obtain the QBD level variable. Then, the phase variable is used to model the dynamics of the underlying environment. Here, we extend the framework of Soewongsono et al. to enable a more general analysis. We develop an efficient recursive algorithm for computing the likelihood of an observed tree under this model and construct several numerical examples to illustrate its application potential. Through our synthetic data examples, we show a range of potential behaviours that could be modelled with this approach. Further, we apply the framework to two empirical examples from comparative phylogenetics (the evolution of range area and body size traits across a phylogeny of 49 mammals) to gain different insights into the evolution of these continuous traits. In this setting duplication of the QBD represents speciation and continuous trait evolution is modelled in a discretised state space. In our empirical examples, we explore the impact of different parameter choices on the corresponding likelihood of observing a given phylogenetic tree and the observed levels at its tips.

q-bio.PE↗

Convergence-divergence models: Generalizations of phylogenetic trees modeling gene flow over time

Phylogenetic trees are simple models of evolutionary processes. They describe conditionally independent divergent evolution of taxa from common ancestors. Phylogenetic trees commonly do not have enough flexibility to adequately model all evolutionary processes. For example, introgressive hybridization, where genes can flow from one taxon to another. Phylogenetic networks model evolution not fully described by a phylogenetic tree. However, many phylogenetic network models assume ancestral taxa merge instantaneously to form ``hybrid'' descendant taxa. In contrast, our convergence-divergence models retain a single underlying ``principal'' tree, but permit gene flow over arbitrary time frames. Alternatively, convergence-divergence models can describe other biological processes leading to taxa becoming more similar over a time frame, such as replicated evolution. Here we present novel maximum likelihood-based algorithms to infer most aspects of $N$-taxon convergence-divergence models, many consistently, using a quartet-based approach. The algorithms can be applied to multiple sequence alignments restricted to genes or genomic windows or to gene presence/absence datasets.

q-bio.PE↗

From trees to traits: A review of advances in PhyloG2P methods and future directions

Mapping genotypes to phenotypes (G2P) is a fundamental goal in biology. So called PhyloG2P methods are a relatively new set of tools that leverage replicated evolution in phylogenetically independent lineages to identify genomic regions associated with traits of interest. Here, we review recent developments in PhyloG2P methods, focusing on three key areas: methods based on replicated amino acid substitutions, methods detecting changes in evolutionary rates, and methods analysing gene duplication and loss. We discuss how the definition and measurement of traits impacts the utility of these methods, arguing that focusing on simple rather than compound traits will lead to more meaningful genotype-phenotype associations. We advocate for the use of methods that work with continuous traits directly rather than collapsing them to binary representations. We examine the strengths and limitations of different approaches to modeling genetic replication, highlighting the importance of explicit modeling of evolutionary processes. Finally, we outline promising future directions, including the integration of population-level variation, as well as epigenetic and environmental information. No one method is likely to identify all genomic regions of interest, so we encourage users to apply multiple methods that are capable of detecting a wide range of associations. The overall aim of this review is to provide practitioners a roadmap for understanding and applying PhyloG2P methods.

q-bio.PE↗

Matrix-analytic methods for the evolution of species trees, gene trees, and their reconciliation

We consider the reconciliation problem, in which the task is to find a mapping of a gene tree into a species tree, so as to maximize the likelihood of such fitting, given the available data. We describe a model for the evolution of the species tree, a subfunctionalisation model for the evolution of the gene tree, and provide an algorithm to compute the likelihood of the reconciliation. We derive our results using the theory of matrix-analytic methods and describe efficient algorithms for the computation of a range of useful metrics. We illustrate the theory with examples and provide the physical interpretations of the discussed quantities, with a focus on the practical applications of the theory to incomplete data.

q-bio.PE↗

Stochastic niche-based models for the evolution of species

There have been many studies to examine whether one trait is correlated with another trait across a group of present-day species (for example, do species with larger brains tend to have longer gestation times. Since the introduction of the phylogenetic comparative method some authors have argued that it is necessary to have a biologically realistic model to generate evolutionary trees that incorporates information about the ecological niche occupied by species. Price presented a simple model along these lines in 1997. He defined a two-dimensional niche space formed by two continuous-valued traits, in which new niches arise with trait values drawn from a bivariate normal distribution. When a new niche arises, it is occupied by a descendant species of whichever current species is closest in ecological niche space. In sequence, more species are then evolved from already-existing species to which they are ecologically closest. Here we explore ways of extending Price's adaptive radiation model. One extension is to increase the dimensionality of the niche space by considering more than two continuous traits. A second extension is to allow both extinction of species (which may leave unoccupied niches) and removal of niches (which causes species occupying them to go extinct). To model this problem, we consider a continuous-time stochastic process which implicitly defines a phylogeny. To explore if trees generated under such a model (or under different parametrizations of the model) are realistic we can compute a variety of summary statistics that can be compared to those of empirically observed phylogenies. For example, there are existing statistics that aim to measure: tree balance, the relative rate of diversification, and phylogenetic signal of traits.

q-bio.PE↗

Modelling gene content across a phylogeny to determine when genes become associated

In this work, we develop a stochastic model of gene gain and loss with the aim of inferring when (if at all) in evolutionary history and association between two genes arises. The data we consider is a species tree along with information on the presence or absence of two genes in each of the species. The biological motivation for our model is that if two genes are involved in the same biochemical pathway, i.e. they are both required for some function, then the rate of gain or loss of one gene in the pathway should depend upon the presence or absence of the other gene in the pathway. However, if the two genes are not functionally linked, then the rate of gain or loss of one gene should be independent of the state of another gene. We simulate data under this model to determine under what conditions a shift from the independent rates class to the dependent rates class can be detected. For example, how large a tree is required and how large a shift in the rates is needed before Akaike information criterion (AIC) supports a model with two rate classes over a simpler model with just one rate class? If a model with two rate classes is preferred, can it correctly detect where on the evolutionary tree the shift occurred?

q-bio.PE↗

The Shape of Phylogenies Under Phase-Type Distributed Times to Speciation and Extinction

Phylogenetic trees are widely used to understand the evolutionary history of organisms. Tree shapes provide information about macroevolutionary processes. However, macroevolutionary models are unreliable for inferring the true processes underlying empirical trees. Here, we propose a flexible and biologically plausible macroevolutionary model for phylogenetic trees where times to speciation or extinction events are drawn from a Coxian phase-type (PH) distribution. First, we show that different choices of parameters in our model lead to a range of tree balances as measured by Aldous' $β$ statistic. In particular, we demonstrate that it is possible to find parameters that correspond well to empirical tree balance. Next, we provide a natural extension of the $β$ statistic to sets of trees. This extension produces less biased estimates of $β$ compared to using the median $β$ values from individual trees. Furthermore, we derive a likelihood expression for the probability of observing any tree with branch lengths under a model with speciation but no extinction. Finally, we illustrate the application of our model by performing both absolute and relative goodness-of-fit tests for two large empirical phylogenies (squamates and angiosperms) that compare models with Coxian PH distributed times to speciation with models that assume exponential or Weibull distributed waiting times. In our numerical analysis, we found that, in most cases, models assuming a Coxian PH distribution provided the best fit.

q-bio.PE↗

The impracticalities of multiplicatively-closed codon models: a retreat to linear alternatives

A matrix Lie algebra is a linear space of matrices closed under the operation $ [A, B] = AB-BA $. The "Lie closure" of a set of matrices is the smallest matrix Lie algebra which contains the set. In the context of Markov chain theory, if a set of rate matrices form a Lie algebra, their corresponding Markov matrices are closed under matrix multiplication; this has been found to be a useful property in phylogenetics. Inspired by previous research involving Lie closures of DNA models, it was hypothesised that finding the Lie closure of a codon model could help to solve the problem of mis-estimation of the non-synonymous/synonymous rate ratio, $ ω$. We propose two different methods of finding a linear space from a model: the first is the \emph{linear closure} which is the smallest linear space which contains the model, and the second is the \emph{linear version} which changes multiplicative constraints in the model to additive ones. For each of these linear spaces we then find the Lie closures of them. Under both methods, it was found that closed codon models would require thousands of parameters, and that any partial solution to this problem that was of a reasonable size violated stochasticity. Investigation of toy models indicated that finding the Lie closure of matrix linear spaces which deviated only slightly from a simple model resulted in a Lie closure that was close to having the maximum number of parameters possible. Given that Lie closures are not practical, we propose further consideration of the two variants of linearly closed models.

q-bio.PE↗

The ancient Operational Code is embedded in the amino acid substitution matrix and aaRS phylogenies

The underlying structure of the canonical amino acid substitution matrix (aaSM) is examined by considering stepwise improvements in the differential recognition of amino acids according to their chemical properties during the branching history of the two aminoacyl-tRNA synthetase (aaRS) superfamilies. The evolutionary expansion of the genetic code is described by a simple parameterization of the aaSM, in which (i) the number of distinguishable amino acid types, (ii) the matrix dimension, and (iii) the number of parameters, each increases by one for each bifurcation in an aaRS phylogeny. Parameterized matrices corresponding to trees in which the size of an amino acid sidechain is the only discernible property behind its categorization as a substrate, exclusively for a Class I or II aaRS, provide a significantly better fit to empirically determined aaSM than trees with random bifurcation patterns. A second split between polar and nonpolar amino acids in each Class effects a vastly greater further improvement. The earliest Class-separated epochs in the phylogenies of the aaRS reflect these enzymes' capability to distinguish tRNAs through the recognition of acceptor stem identity elements via the minor (Class I) and major (Class II) helical grooves, which is how the ancient Operational Code functioned. The advent of tRNA recognition using the anticodon loop supports the evolution of the optimal map of amino acid chemistry found in the later Genetic Code, an essentially digital categorization, in which polarity is the major functional property, compensating for the unrefined, haphazard differentiation of amino acids achieved by the Operational Code.

q-bio.PE↗

Exploring the consequences of lack of closure in codon models

Models of codon evolution are commonly used to identify positive selection. Positive selection is typically a heterogeneous process, i.e., it acts on some branches of the evolutionary tree and not others. Previous work on DNA models showed that when evolution occurs under a heterogeneous process it is important to consider the property of model closure, because non-closed models can give biased estimates of evolutionary processes. The existing codon models that account for the genetic code are not closed; to establish this it is enough to show that they are not linear (meaning that the sum of two codon rate matrices in the model is not a matrix in the model). This raises the concern that a single codon model fit to a heterogeneous process might mis-estimate both the effect of selection and branch lengths. Codon models are typically constructed by choosing an underlying DNA model (e.g., HKY) that acts identically and independently at each codon position, and then applying the genetic code via the parameter $ω$ to modify the rate of transitions between codons that code for different amino acids. Here we use simulation to investigate the accuracy of estimation of both the selection parameter $ω$ and branch lengths in cases where the underlying DNA process is heterogeneous but $ω$ is constant. We find that both $ω$ and branch lengths can be mis-estimated in these scenarios. Errors in $ω$ were usually less than 2% but could be as high as 17%. We also assessed if choosing different underlying DNA models had any affect on accuracy, in particular we assessed if using closed DNA models gave any advantage. However, a DNA model being closed does not imply that the codon model constructed from it is closed, and in general we found that using closed DNA models did not decrease errors in the estimation of $ω$.

q-bio.PE↗

Distinguishing between convergent evolution and violation of the molecular clock

We give a non-technical introduction to convergence-divergence models, a new modeling approach for phylogenetic data that allows for the usual divergence of species post speciation but also allows for species to converge, i.e. become more similar over time. By examining the $3$-taxon case in some detail we illustrate that phylogeneticists have been "spoiled" in the sense of not having to think about the structural parameters in their models by virtue of the strong assumption that evolution is treelike. We show that there are not always good statistical reasons to prefer the usual class of treelike models over more general convergence-divergence models. Specifically we show many $3$-taxon datasets can be equally well explained by supposing violation of the molecular clock due to change in the rate of evolution along different edges, or by keeping the assumption of a constant rate of evolution but instead assuming that evolution is not a purely divergent process. Given the abundance of evidence that evolution is not strictly treelike, our discussion is an illustration that as phylogeneticists we often need to think clearly about the structural form of the models we use.

q-bio.PE↗

Maximum likelihood estimates of pairwise rearrangement distances

Accurate estimation of evolutionary distances between taxa is important for many phylogenetic reconstruction methods. In the case of bacteria, distances can be estimated using a range of different evolutionary models, from single nucleotide polymorphisms to large-scale genome rearrangements. In the case of sequence evolution models (such as the Jukes-Cantor model and associated metric) have been used to correct pairwise distances. Similar correction methods for genome rearrangement processes are required to improve inference. Current attempts at correction fall into 3 categories: Empirical computational studies, Bayesian/MCMC approaches, and combinatorial approaches. Here we introduce a maximum likelihood estimator for the inversion distance between a pair of genomes, using the group-theoretic approach to modelling inversions introduced recently. This MLE functions as a corrected distance: in particular, we show that because of the way sequences of inversions interact with each other, it is quite possible for minimal distance and MLE distance to differently order the distances of two genomes from a third. This has obvious implications for the use of minimal distance in phylogeny reconstruction. The work also tackles the above problem allowing free rotation of the genome. Generally a frame of reference is locked, and all computation made accordingly. This work incorporates the action of the dihedral group so that distance estimates are free from any a priori frame of reference.

q-bio.PE↗

Comparison of three Statistical Classification Techniques for Maser Identification

We applied three statistical classification techniques - linear discriminant analysis (LDA), logistic regression and random forests - to three astronomical datasets associated with searches for interstellar masers. We compared the performance of these methods in identifying whether specific mid-infrared or millimetre continuum sources are likely to have associated interstellar masers. We also discuss the ease, or otherwise, with which the results of each classification technique can be interpreted. Non-parametric methods have the potential to make accurate predictions when there are complex relationships between critical parameters. We found that for the small datasets the parametric methods logistic regression and LDA performed best, for the largest dataset the non-parametric method of random forests performed with comparable accuracy to parametric techniques, rather than any significant improvement. This suggests that at least for the specific examples investigated here accuracy of the predictions obtained is not being limited by the use of parametric models. We also found that for LDA, transformation of the data to match a normal distribution in the input parameters led to big improvements in accuracy. The different classification techniques had significant overlap in their predictions, further astronomical observations will enable the accuracy of these predictions to be tested.

astro-ph.IM↗

A tensorial approach to the inversion of group-based phylogenetic models

Using a tensorial approach, we show how to construct a one-one correspondence between pattern probabilities and edge parameters for any group-based model. This is a generalisation of the "Hadamard conjugation" and is equivalent to standard results that use Fourier analysis. In our derivation we focus on the connections to group representation theory and emphasize that the inversion is possible because, under their usual definition, group-based models are defined for abelian groups only. We also argue that our approach is elementary in the sense that it can be understood as simple matrix multiplication where matrices are rectangular and indexed by ordered-partitions of varying sizes.

q-bio.PE↗

Low-parameter phylogenetic estimation under the general Markov model

In their 2008 and 2009 papers, Sumner and colleagues introduced the "squangles" - a small set of Markov invariants for phylogenetic quartets. The squangles are consistent with the general Markov model (GM) and can be used to infer quartets without the need to explicitly estimate all parameters. As GM is inhomogeneous and hence non-stationary, the squangles are expected to perform well compared to standard approaches when there are changes in base-composition amongst species. However, GM includes the IID assumption, so the squangles should be confounded by data generated with invariant sites or with rate-variation across sites. Here we implement the squangles in a least-squares setting that returns quartets weighted by either confidence or internal edge lengths; and use these as input into a variety of quartet-based supertree methods. For the first time, we quantitatively investigate the robustness of the squangles to the breaking of IID assumptions on both simulated and real data sets; and we suggest a modification that improves the performance of the squangles in the presence of invariant sites. Our conclusion is that the squangles provide a novel tool for phylogenetic estimation that is complementary to methods that explicitly account for rate-variation across sites, but rely on homogeneous - and hence stationary - models.

q-bio.QM↗

Novel Distances for Dollo Data

We investigate distances on binary (presence/absence) data in the context of a Dollo process, where a trait can only arise once on a phylogenetic tree but may be lost many times. We introduce a novel distance, the Additive Dollo Distance (ADD), which is consistent for data generated under a Dollo model, and show that it has some useful theoretical properties including an intriguing link to the LogDet distance. Simulations of Dollo data are used to compare a number of binary distances including ADD, LogDet, Nei Li and some simple, but to our knowledge previously unstudied, variations on common binary distances. The simulations suggest that ADD outperforms other distances on Dollo data. Interestingly, we found that the LogDet distance performs poorly in the context of a Dollo process, which may have implications for its use in connection with conditioned genome reconstruction. We apply the ADD to two Diversity Arrays Technology (DArT) datasets, one that broadly covers Eucalyptus species and one that focuses on the Eucalyptus series Adnataria. We also reanalyse gene family presence/absence data on bacteria from the COG database and compare the results to previous phylogenies estimated using the conditioned genome reconstruction approach.

q-bio.QM↗