SearcharxivSearch

arXiv subjects

Mevin B. Hooten

Publications and source records attributed to Mevin B. Hooten.

At least 19 recordsLinked to original sources

Recursive Adaptive Importance Sampling with Optimal Replenishment

Increased access to computing resources has led to the development of algorithms that can run efficiently on multi-core processing units or in distributed computing environments. In the context of Bayesian inference, many parallel computing approaches to fit statistical models have been proposed using Markov chain Monte Carlo methods, but they either have limited gains due to high latency cost or rely on model-specific decompositions. Alternatively, adaptive importance sampling, sequential Monte Carlo, and recursive Bayesian methods provide a parallel-friendly and asymptotically exact framework with well-developed theory for error estimation. We consider a recursive adaptive importance sampling approach that alternates between fast recursive weight updates and sample replenishment steps to balance computational effort while ensuring sample quality. A naive implementation can be unstable and inefficient, so we derive theoretical results to determine the optimal replenishment frequency based on the posterior concentration rate, thus attaining an asymptotically optimal trade-off. We demonstrate the efficacy of our method in simulated experiments and an application of sea surface temperature prediction in the Gulf of Mexico using Gaussian processes.

stat.ME

Making Recursive Bayesian Inference Robust

While Bayesian inference has become increasingly popular with advances in computational resources, its algorithms can be computationally prohibitive and may not scale with large datasets. This has led to growing interest in alternative algorithms, such as approximation methods and variants of Markov chain Monte Carlo. Among these approaches, prior proposal-recursive Bayesian (PP-RB) inference facilitates scalable Bayesian computation by recursively updating the posterior distribution across stages and utilizing parallel computing resources. While the well-known ``degeneracy'' issue in PP-RB has been studied, another limitation that PP-RB can yield incorrect inferences when posterior distributions shift substantially between stages has remained unsolved. To address this, we propose parallel-tempered prior proposal-recursive Bayesian (PPP-RB) inference, which extends PP-RB by leveraging the key idea underlying Metropolis-coupled Markov chain Monte Carlo. We show both theoretically and empirically that PPP-RB targets the true posterior distribution. We illustrate PPP-RB through numerical studies and real data analysis in application to earthquake count data and sea surface salinity in the North Atlantic region. In these applications, we compare PPP-RB with PP-RB and a standard MCMC, demonstrating that PPP-RB is more efficient in terms of effective sample size per elapsed time.

stat.ME

A multi-stage Bayesian approach to fit spatial point process models

Spatial point process (SPP) models are commonly used to analyze point pattern data in many fields, including presence-only data in ecology. Existing exact Bayesian methods for fitting these models are computationally expensive because they require approximating an intractable integral each time parameters are updated and often involve algorithm supervision (i.e., tuning in the Bayesian setting). We propose a flexible, efficient, and exact multi-stage recursive Bayesian approach to fitting SPP models that leverages parallel computing resources to obtain realizations from the joint posterior, which can then be used to obtain inference on derived quantities. We outline potential extensions, including a framework for analyzing study designs with compact observation windows and a neural network basis expansion for increased model flexibility. We demonstrate this approach and its extensions using a simulation study and analyze data from aerial imagery surveys to improve our understanding of spatially explicit abundance of harbor seal (Phoca vitulina) pups in Johns Hopkins Inlet, a protected tidewater glacial fjord in Glacier Bay National Park, Alaska.

stat.ME

Dyadic Flow Models for Nonstationary Gene Flow in Landscape Genomics

The field of landscape genomics aims to infer how landscape features affect gene flow across space. Most landscape genomic frameworks assume the isolation-by-distance and isolation-by-resistance hypotheses, which propose that genetic dissimilarity increases as a function of distance and as a function of cumulative landscape resistance, respectively. While these hypotheses are valid in certain settings, other mechanisms may affect gene flow. For example, the gene flow of invasive species may depend on founder effects and multiple introductions. Such mechanisms are not considered in modern landscape genomic models. We extend dyadic models to allow for mechanisms that range-shifting and/or invasive species may experience by introducing dyadic spatially-varying coefficients (DSVCs) defined on source-destination pairs. The DSVCs allow the effects of landscape on gene flow to vary across space, capturing nonstationary and asymmetric connectivity. Additionally, we incorporate explicit landscape features as connectivity covariates, which are localized to specific regions of the spatial domain and may function as barriers or corridors to gene flow. Such covariates are central to colonization and invasion, where spread accelerates along corridors and slows across landscape barriers. The proposed framework accommodates colonization-specific processes while retaining the ability to assess landscape influences on gene flow. Our case study of the highly invasive cheatgrass (Bromus tectorum) demonstrates the necessity of accounting for nonstationarity gene flow in range-shifting species.

stat.AP

Spatial Hyperspheric Models for Compositional Data

Compositional observations are an increasingly prevalent data source in spatial statistics. Analysis of such data is typically done on log-ratio transformations or via Dirichlet regression. However, these approaches often make unnecessarily strong assumptions (e.g., strictly positive components, exclusively negative correlations). An alternative approach uses square-root transformed compositions and directional distributions. Such distributions naturally allow for zero-valued components and positive correlations, yet they may include support outside the non-negative orthant and are not generative for compositional data. To overcome this challenge, we truncate the elliptically symmetric angular Gaussian (ESAG) distribution to the non-negative orthant. Additionally, we propose a spatial hyperspheric regression model that contains fixed and random multivariate spatial effects. The proposed model also contains a term that can be used to propagate uncertainty that may arise from precursory stochastic models (i.e., machine learning classification). We used our model in a simulation study and for a spatial analysis of classified bioacoustic signals of the Dryobates pubescens (downy woodpecker).

stat.ME

A Unified Bayesian Framework for Modeling Measurement Error in Multinomial Data

Measurement error in multinomial data is a well-known and well-studied inferential problem that is encountered in many fields, including engineering, biomedical and omics research, ecology, finance, official statistics, and social sciences. Methods developed to accommodate measurement error in multinomial data are typically equipped to handle false negatives or false positives, but not both. We provide a unified framework for accommodating both forms of measurement error using a Bayesian hierarchical approach. We demonstrate the proposed method's performance on simulated data and apply it to acoustic bat monitoring and official crime data.

stat.ME

Spatial Knockoff Bayesian Variable Selection in Genome-Wide Association Studies

High-dimensional variable selection has emerged as one of the prevailing statistical challenges in the big data revolution. Many variable selection methods have been adapted for identifying single nucleotide polymorphisms (SNPs) linked to phenotypic variation in genome-wide association studies. We develop a Bayesian variable selection regression model for identifying SNPs linked to phenotypic variation. We modify our Bayesian variable selection regression models to control the false discovery rate of SNPs using a knockoff variable approach. We reduce spurious associations by regressing the phenotype of interest against a set of basis functions that account for the relatedness of individuals. Using a restricted regression approach, we simultaneously estimate the SNP-level effects while removing variation in the phenotype that can be explained by population structure. We also accommodate the spatial structure among causal SNPs by modeling their inclusion probabilities jointly with a reduced rank Gaussian process. In a simulation study, we demonstrate that our spatial Bayesian variable selection regression model controls the false discovery rate and increases power when the relevant SNPs are clustered. We conclude with an analysis of Arabidopsis thaliana flowering time, a polygenic trait that is confounded with population structure, and find the discoveries of our method cluster near described flowering time genes.

stat.ME

Dynamic Population Models with Temporal Preferential Sampling to Infer Phenology

To study population dynamics, ecologists and wildlife biologists use relative abundance data, which are often subject to temporal preferential sampling. Temporal preferential sampling occurs when sampling effort varies across time. To account for preferential sampling, we specify a Bayesian hierarchical abundance model that considers the dependence between observation times and the ecological process of interest. The proposed model improves abundance estimates during periods of infrequent observation and accounts for temporal preferential sampling in discrete time. Additionally, our model facilitates posterior inference for population growth rates and mechanistic phenometrics. We apply our model to analyze both simulated data and mosquito count data collected by the National Ecological Observatory Network. In the second case study, we characterize the population growth rate and abundance of several mosquito species in the Aedes genus.

stat.ME

Source Reconstruction for Spatio-Temporal Physical Statistical Models

In many applications, a signal is deformed by well-understood dynamics before it can be measured. For example, when a pollutant enters a river, it immediately begins dispersing, flowing, settling, and reacting. If the pollutant enters at a single point, its concentration can be measured before it enters the complex dynamics of the river system. However, in the case of a non-point source pollutant, it is not clear how to efficiently measure its source. One possibility is to record concentration measurements in the river, but this signal is masked by the fluid dynamics of the river. Specifically, concentration is governed by the advection-diffusion-reaction PDE, with an unknown source term. We propose a method to statistically reconstruct a source term from these PDE-deformed measurements. Our method is general and applies to any linear PDE. This method has important applications in the study of environmental DNA and non-point source pollution.

stat.OT

Latent trajectory models for spatio-temporal dynamics in Alaskan ecosystems

The Alaskan landscape has undergone substantial changes in recent decades, most notably the expansion of shrubs and trees across the Arctic. We developed a dynamic statistical model to quantify the impact of climate change on the structural transformation of ecosystems using remotely sensed imagery. We used latent trajectory processes in a hierarchical framework to model dynamic state probabilities that evolve annually, from which we derived transition probabilities between ecotypes. Our latent trajectory model accommodates temporal irregularity in survey intervals and uses spatio-temporally heterogeneous climate drivers to infer rates of land cover transitions. We characterized multi-scale spatial correlation induced by plot and subplot arrangement in our study system. We also developed a Polya-Gamma sampling strategy to improve computation. Our model facilitates inference on the response of ecosystems to shifts in the climate and can be used to predict future land cover transitions under various climate scenarios.

stat.AP

Bayesian Inverse Reinforcement Learning for Collective Animal Movement

Agent-based methods allow for defining simple rules that generate complex group behaviors. The governing rules of such models are typically set a priori and parameters are tuned from observed behavior trajectories. Instead of making simplifying assumptions across all anticipated scenarios, inverse reinforcement learning provides inference on the short-term (local) rules governing long term behavior policies by using properties of a Markov decision process. We use the computationally efficient linearly-solvable Markov decision process to learn the local rules governing collective movement for a simulation of the self propelled-particle (SPP) model and a data application for a captive guppy population. The estimation of the behavioral decision costs is done in a Bayesian framework with basis function smoothing. We recover the true costs in the SPP simulation and find the guppies value collective movement more than targeted movement toward shelter.

cs.LG

Greater Than the Sum of its Parts: Computationally Flexible Bayesian Hierarchical Modeling

We propose a multistage method for making inference at all levels of a Bayesian hierarchical model (BHM) using natural data partitions to increase efficiency by allowing computations to take place in parallel form using software that is most appropriate for each data partition. The full hierarchical model is then approximated by the product of independent normal distributions for the data component of the model. In the second stage, the Bayesian maximum {\it a posteriori} (MAP) estimator is found by maximizing the approximated posterior density with respect to the parameters. If the parameters of the model can be represented as normally distributed random effects then the second stage optimization is equivalent to fitting a multivariate normal linear mixed model. This method can be extended to account for common fixed parameters shared between data partitions, as well as parameters that are distinct between partitions. In the case of distinct parameter estimation, we consider a third stage that re-estimates the distinct parameters for each data partition based on the results of the second stage. This allows more information from the entire data set to properly inform the posterior distributions of the distinct parameters. The method is demonstrated with two ecological data sets and models, a random effects GLM and an Integrated Population Model (IPM). The multistage results were compared to estimates from models fit in single stages to the entire data set. Both examples demonstrate that multistage point and posterior standard deviation estimates closely approximate those obtained from fitting the models with all data simultaneously and can therefore be considered for fitting hierarchical Bayesian models when it is computationally prohibitive to do so in one step.

stat.ME

Accounting for phenology in the analysis of animal movement

The analysis of animal tracking data provides an important source of scientific understanding and discovery in ecology. Observations of animal trajectories using telemetry devices provide researchers with information about the way animals interact with their environment and each other. For many species, specific geographical features in the landscape can have a strong effect on behavior. Such features may correspond to a single point (e.g., dens or kill sites), or to higher-dimensional subspaces (e.g., rivers or lakes). Features may be relatively static in time (e.g., coastlines or home-range centers), or may be dynamic (e.g., sea ice extent or areas of high-quality forage for herbivores). We introduce a novel model for animal movement that incorporates active selection for dynamic features in a landscape. Our approach is motivated by the study of polar bear (Ursus maritimus) movement. During the sea ice melt season, polar bears spend much of their time on sea ice above shallow, biologically productive water where they hunt seals. The changing distribution and characteristics of sea ice throughout the late spring through early fall means that the location of valuable habitat is constantly shifting. We develop a model for the movement of polar bears that accounts for the effect of this important landscape feature. We introduce a two-stage procedure for approximate Bayesian inference that allows us to analyze over 300,000 observed locations of 186 polar bears from 2012--2016. We use our proposed model to answer a particular question posed by wildlife managers who seek to cluster polar bears from the Beaufort and Chukchi seas into sub-populations.

stat.ME

Running on empty: Recharge dynamics from animal movement data

Vital rates such as survival and recruitment have always been important in the study of population and community ecology. At the individual level, physiological processes such as energetics are critical in understanding biomechanics and movement ecology and also scale up to influence food webs and trophic cascades. Although vital rates and population-level characteristics are tied with individual-level animal movement, most statistical models for telemetry data are not equipped to provide inference about these relationships because they lack the explicit, mechanistic connection to physiological dynamics. We present a framework for modeling telemetry data that explicitly includes an aggregated physiological process associated with decision making and movement in heterogeneous environments. Our framework accommodates a wide range of movement and physiological process specifications. We illustrate a specific model formulation in continuous-time to provide direct inference about gains and losses associated with physiological processes based on movement. Our approach can also be extended to accommodate auxiliary data when available. We demonstrate our model to infer mountain lion (in Colorado, USA) and African buffalo (in Kruger National Park, South Africa) recharge dynamics.

stat.AP

Animal Movement Models with Mechanistic Selection Functions

A suite of statistical methods are used to study animal movement. Most of these methods treat animal telemetry data in one of three ways: as discrete processes, as continuous processes, or as point processes. We briefly review each of these approaches and then focus in on the latter. In the context of point processes, so-called resource selection analyses are among the most common way to statistically treat animal telemetry data. However, most resource selection analyses provide inference based on approximations of point process models. The forms of these models have been limited to a few types of specifications that provide inference about relative resource use and, less commonly, probability of use. For more general spatio-temporal point process models, the most common type of analysis often proceeds with a data augmentation approach that is used to create a binary data set that can be analyzed with conditional logistic regression. We show that the conditional logistic regression likelihood can be generalized to accommodate a variety of alternative specifications related to resource selection. We then provide an example of a case where a spatio-temporal point process model coincides with that implied by a mechanistic model for movement expressed as a partial differential equation derived from first principles of movement. We demonstrate that inference from this form of point process model is intuitive (and could be useful for management and conservation) by analyzing a set of telemetry data from a mountain lion in Colorado, USA, to understand the effects of spatially explicit environmental conditions on movement behavior of this species.

stat.ME

Hierarchical approaches for flexible and interpretable binary regression models

Binary regression models are ubiquitous in virtually every scientific field. Frequently, traditional generalized linear models fail to capture the variability in the probability surface that gives rise to the binary observations and novel methodology is required. This has generated a substantial literature comprised of binary regression models motivated by various applications. We describe a novel organization of generalizations to traditional binary regression methods based on the familiar three-part structure of generalized linear models (random component, systematic component, link function). This new perspective facilitates both the comparison of existing approaches, and the development of novel, flexible models with interpretable parameters that capture application-specific data generating mechanisms. We use our proposed organizational structure to discuss some concerns with certain existing models for binary data based on quantile regression. We then use the framework to develop several new binary regression models tailored to occupancy data for European red squirrels (Sciurus vulgaris).

stat.ME

Making Recursive Bayesian Inference Accessible

Bayesian models provide recursive inference naturally because they can formally reconcile new data and existing scientific information. However, popular use of Bayesian methods often avoids priors that are based on exact posterior distributions resulting from former studies. Two existing Recursive Bayesian methods are: Prior- and Proposal-Recursive Bayes. Prior-Recursive Bayes uses Bayesian updating, fitting models to partitions of data sequentially, and provides a way to accommodate new data as they become available using the posterior from the previous stage as the prior in the new stage based on the latest data. Proposal-Recursive Bayes is intended for use with hierarchical Bayesian models and uses a set of transient priors in first stage independent analyses of the data partitions. The second stage of Proposal-Recursive Bayes uses the posteriors from the first stage as proposals in an MCMC algorithm to fit the full model. We combine Prior- and Proposal-Recursive concepts to fit any Bayesian model, and often with computational improvements. We demonstrate our method with two case studies. Our approach has implications for big data, streaming data, and optimal adaptive design situations.

stat.ME

Predicting paleoclimate from compositional data using multivariate Gaussian process inverse prediction

Multivariate compositional count data arise in many applications including ecology, microbiology, genetics, and paleoclimate. A frequent question in the analysis of multivariate compositional count data is what values of a covariate(s) give rise to the observed composition. Learning the relationship between covariates and the compositional count allows for inverse prediction of unobserved covariates given compositional count observations. Gaussian processes provide a flexible framework for modeling functional responses with respect to a covariate without assuming a functional form. Many scientific disciplines use Gaussian process approximations to improve prediction and make inference on latent processes and parameters. When prediction is desired on unobserved covariates given realizations of the response variable, this is called inverse prediction. Because inverse prediction is mathematically and computationally challenging, predicting unobserved covariates often requires fitting models that are different from the hypothesized generative model. We present a novel computational framework that allows for efficient inverse prediction using a Gaussian process approximation to generative models. Our framework enables scientific learning about how the latent processes co-vary with respect to covariates while simultaneously providing predictions of missing covariates. The proposed framework is capable of efficiently exploring the high dimensional, multi-modal latent spaces that arise in the inverse problem. To demonstrate flexibility, we apply our method in a generalized linear model framework to predict latent climate states given multivariate count data. Based on cross-validation, our model has predictive skill competitive with current methods while simultaneously providing formal, statistical inference on the underlying community dynamics of the biological system previously not available.

stat.ME