SearcharxivSearch

arXiv subjects

Samuel L. Hong

Publications and source records attributed to Samuel L. Hong.

3 recordsLinked to original sources

An adaptive time-tree transition kernel for Bayesian phylogenetic inference

Bayesian phylogenetic and phylodynamic analyses can be very time-consuming, owing to the combination of complex models that are used to estimate key parameters from increasingly large genomic data sets and their associated metadata. The use of high-performance computer hardware can -- to a certain extent -- alleviate the computational burden and markedly decrease the time to results. Still, even converging to the posterior can be a lengthy endeavour, with the burn-in aspect of such analyses potentially taking days or even weeks for large data sets. One of the key aspects that hampers performance in Bayesian phylogenetic inference is the efficiency with which tree topology proposals explore tree space. We here propose a novel adaptive tree transition kernel, which we call `subTreeLeap' (STL), which involves modifying the phylogeny by walking along patristic distance paths in the tree according to an adaptable radius parameter. STL is a general proposal, which can be used with contemporaneous or time-calibrated sequence data, being particularly suited to the latter due to respecting temporal precedence constraints. We carefully assess its impact on convergence and statistical mixing of the exploration of posterior tree space, by comparison to replicate ``golden runs'' obtained from lengthy analyses of empirical data under standard tree transition kernels. We find that STL successfully explores the same posterior tree space as standard kernels, but often does so in a more efficient manner. We discuss limitations as well as future potential improvements to STL that could substantially increase the speed at which Bayesian phylogenetic inferences are obtained.

q-bio.PE

Detecting Evolutionary Change-Points with Branch-Specific Substitution Models and Shrinkage Priors

Branch-specific substitution models are popular for detecting evolutionary change-points, such as shifts in selective pressure. However, applying such models typically requires prior knowledge of change-point locations on the phylogeny or faces scalability issues with large data sets. To address both limitations, we integrate branch-specific substitution models with shrinkage priors to automatically identify change-points without prior knowledge, while simultaneously estimating distinct substitution parameters for each branch. To enable tractable inference under this high-dimensional model, we develop an analytical gradient algorithm for the branch-specific substitution parameters where the computational time is linear in the number of parameters. We apply this gradient algorithm to infer selection pressure dynamics in the evolution of the BRCA1 gene in primates and mutational dynamics in viral sequences from the recent mpox epidemic. Our novel algorithm enhances inference efficiency, achieving up to a 126-fold speedup per iteration in maximum likelihood optimization when compared to central difference numerical gradient method and up to a 2026-fold improvement in computational performance within a Bayesian framework using Hamiltonian Monte Carlo sampler compared to conventional univariate random walk sampler.

q-bio.PE

On the importance of assessing topological convergence in Bayesian phylogenetic inference

Modern phylogenetics research is often performed within a Bayesian framework, using sampling algorithms such as Markov chain Monte Carlo (MCMC) to approximate the posterior distribution. These algorithms require careful evaluation of the quality of the generated samples. Within the field of phylogenetics, one frequently adopted diagnostic approach is to evaluate the effective sample size (ESS) and to investigate trace graphs of the sampled parameters. A major limitation of these approaches is that they are developed for continuous parameters and therefore incompatible with a crucial parameter in these inferences: the tree topology. Several recent advancements have aimed at extending these diagnostics to topological space. In this reflection paper, we present two case studies - one on Ebola virus and one on HIV - illustrating how these topological diagnostics can contain information not found in standard diagnostics, and how decisions regarding which of these diagnostics to compute can impact inferences regarding MCMC convergence and mixing. Our results show the importance of running multiple replicate analyses and of carefully assessing topological convergence using the output of these replicate analyses. To this end, we illustrate different ways of assessing and visualizing the topological convergence of these replicates. Given the major importance of detecting convergence and mixing issues in Bayesian phylogenetic analyses, the lack of a unified approach to this problem warrants further action, especially now that additional tools are becoming available to researchers.

q-bio.PE