SearcharxivSearch

arXiv subjects

Evan Sidrow

Publications and source records attributed to Evan Sidrow.

5 recordsLinked to original sources

Variational phylogenetic inference with products over bipartitions

Bayesian phylogenetics is vital for understanding evolutionary dynamics, and requires accurate and efficient approximation of posterior distributions over trees. In this work, we develop a variational Bayesian approach for ultrametric phylogenetic trees. We present a novel variational family based on coalescent times of a single-linkage clustering and derive a closed-form density for the resulting distribution over trees. Unlike existing methods for ultrametric trees, our method performs inference over all of tree space, it does not require any Markov chain Monte Carlo subroutines, and our variational family is differentiable. Through experiments on benchmark genomic datasets and an application to the viral RNA of SARS-CoV-2, we demonstrate that our method achieves competitive accuracy while requiring significantly fewer gradient evaluations than existing state-of-the-art techniques.

stat.ML

Incorporating sparse labels into hidden Markov models using weighted likelihoods improves accuracy and interpretability in biologging studies

Ecologists often use a hidden Markov model to decode a latent process, such as a sequence of an animal's behaviours, from an observed biologging time series. Modern technological devices such as video recorders and drones now allow researchers to directly observe an animal's behaviour. Using these observations as labels of the latent process can improve a hidden Markov model's accuracy when decoding the latent process. However, many wild animals are observed infrequently. Including such rare labels often has a negligible influence on parameter estimates, which in turn does not meaningfully improve the accuracy of the decoded latent process. We introduce a weighted likelihood approach that increases the relative influence of labelled observations. We use this approach to develop two hidden Markov models to decode the foraging behaviour of killer whales (Orcinus orca) off the coast of British Columbia, Canada. Using cross-validated evaluation metrics, we show that our weighted likelihood approach produces more accurate and understandable decoded latent processes compared to existing methods. Thus, our method effectively leverages sparse labels to enhance researchers' ability to accurately decode hidden processes across various fields.

stat.ME

An introduction to statistical models used to characterize species-habitat associations with animal movement data

Understanding species-habitat associations is fundamental to ecological sciences and for species conservation. Consequently, various statistical approaches have been designed to infer species-habitat associations. Due to their conceptual and mathematical differences, these methods can yield contrasting results. We describe and compare commonly used statistical models that relate animal movement data to environmental data, including resource selection functions (RSF), step-selection functions (SSF), and hidden Markov models (HMMs). We demonstrate differences in assumptions and highlighting advantages and limitations of each method. Additionally, we provide guidance on selecting the most appropriate statistical method based on the scale of the data and intended inference. To illustrate the varying ecological insights derived from each model, we apply them to the movement track of a single ringed seal in a case study. We demonstrate that each model yields varying ecological insights. For example, while the selection coefficient values from RSFs appear to show a stronger positive relationship with prey diversity than those of the SSFs, when we accounted for the autocorrelation in the data none of these relationships with prey diversity were statistically significant. The HMM reveals variable associations with prey diversity across different behaviors. Notably, the three models identified different important areas. This case study highlights the critical significance of selecting the appropriate model as an essential step in the process of identifying species-habitat relationships and specific areas of importance. Our review provides the foundational information required for making informed decisions when choosing the most suitable statistical methods to address specific questions, such as identifying protected zones, understanding movement patterns, or studying behaviours.

stat.ME

Variance-Reduced Stochastic Optimization for Efficient Inference of Hidden Markov Models

Hidden Markov models (HMMs) are popular models to identify a finite number of latent states from sequential data. However, fitting them to large data sets can be computationally demanding because most likelihood maximization techniques require iterating through the entire underlying data set for every parameter update. We propose a novel optimization algorithm that updates the parameters of an HMM without iterating through the entire data set. Namely, we combine a partial E step with variance-reduced stochastic optimization within the M step. We prove the algorithm converges under certain regularity conditions. We test our algorithm empirically using a simulation study as well as a case study of kinematic data collected using suction-cup attached biologgers from eight northern resident killer whales (Orcinus orca) off the western coast of Canada. In both, our algorithm converges in fewer epochs and to regions of higher likelihood compared to standard numerical optimization techniques. Our algorithm allows practitioners to fit complicated HMMs to large time-series data sets more efficiently than existing baselines.

stat.CO

Modelling multi-scale state-switching functional data with hidden Markov models

Data sets comprised of sequences of curves sampled at high frequencies in time are increasingly common in practice, but they can exhibit complicated dependence structures that cannot be modelled using common methods of Functional Data Analysis (FDA). We detail a hierarchical approach which treats the curves as observations from a hidden Markov model (HMM). The distribution of each curve is then defined by another fine-scale model which may involve auto-regression and require data transformations using moving-window summary statistics or Fourier analysis. This approach is broadly applicable to sequences of curves exhibiting intricate dependence structures. As a case study, we use this framework to model the fine-scale kinematic movement of a northern resident killer whale (Orcinus orca) off the coast of British Columbia, Canada. Through simulations, we show that our model produces more interpretable state estimation and more accurate parameter estimates compared to existing methods.

stat.ME