Searcharxiv⌕ Search

arXiv subjects

Trevor J. Hefley

Publications and source records attributed to Trevor J. Hefley.

7 recordsLinked to original sources

Data fusion of distance sampling and capture-recapture data

Species distribution models (SDMs) are increasingly used in ecology, biogeography, and wildlife management to learn about the species-habitat relationships and abundance across space and time. Distance sampling (DS) and capture-recapture (CR) are two widely collected data types to learn about species-habitat relationships and abundance; still, they are seldomly used in SDMs due to the lack of spatial coverage. However, data fusion of the two data sources can increase spatial coverage, which can reduce parameter uncertainty and make predictions more accurate, and therefore, can be used for species distribution modeling. We developed a model-based approach for data fusion of DS and CR data. Our modeling approach accounts for two common missing data issues: 1) missing individuals that are missing not at random (MNAR) and 2) partially missing location information. Using a simulation experiment, we evaluated the performance of our modeling approach and compared it to existing approaches that use ad-hoc methods to account for missing data issues. Our results show that our approach provides unbiased parameter estimates with increased efficiency compared to the existing approaches. We demonstrated our approach using data collected for Grasshopper Sparrows (Ammodramus savannarum) in north-eastern Kansas, USA.

stat.ME↗

Recovering individual-level spatial inference from aggregated binary data

Binary regression models are commonly used in disciplines such as epidemiology and ecology to determine how spatial covariates influence individuals. In many studies, binary data are shared in a spatially aggregated form to protect privacy. For example, rather than reporting the location and result for each individual that was tested for a disease, researchers may report that a disease was detected or not detected within geopolitical units. Often, the spatial aggregation process obscures the values of response variables, spatial covariates, and locations of each individual, which makes recovering individual-level inference difficult. We show that applying a series of transformations, including a change of support, to a bivariate point process model allows researchers to recover individual-level inference for spatial covariates from spatially aggregated binary data. The series of transformations preserves the convenient interpretation of desirable binary regression models that are commonly applied to individual-level data. Using a simulation experiment, we compare the performance of our proposed method under varying types of spatial aggregation against the performance of standard approaches using the original individual-level data. We illustrate our method by modeling individual-level probability of infection using a data set that has been aggregated to protect an at-risk and endangered species of bats. Our simulation experiment and data illustration demonstrate the utility of the proposed method when access to original non-aggregated data is impractical or prohibited.

stat.ME↗

Using machine learning to identify nontraditional spatial dependence in occupancy data

Spatial models for occupancy data are used to estimate and map the true presence of a species, which may depend on biotic and abiotic factors as well as spatial autocorrelation. Traditionally researchers have accounted for spatial autocorrelation in occupancy data by using a correlated normally distributed site-level random effect, which might be incapable of identifying nontraditional spatial dependence such as discontinuities and abrupt transitions. Machine learning approaches have the potential to identify and model nontraditional spatial dependence, but these approaches do not account for observer errors such as false absences. By combining the flexibility of Bayesian hierarchal modeling and machine learning approaches, we present a general framework to model occupancy data that accounts for both traditional and nontraditional spatial dependence as well as false absences. We demonstrate our framework using six synthetic occupancy data sets and two real data sets. Our results demonstrate how to identify and model both traditional and nontraditional spatial dependence in occupancy data which enables a broader class of spatial occupancy models that can be used to improve predictive accuracy and model adequacy.

stat.AP↗

When and where: estimating the date and location of introduction for exotic pests and pathogens

A fundamental question during the outbreak of a novel disease or invasion of an exotic pest is: At what location and date was it first introduced? With this information, future introductions can be anticipated and perhaps avoided. Point process models are commonly used for mapping species distribution and disease occurrence. If the time and location of introductions were known, then point process models could be used to map and understand the factors that influence introductions; however, rarely is the process of introduction directly observed. We propose embedding a point process within hierarchical Bayesian models commonly used to understand the spatio-temporal dynamics of invasion. Including a point process within a hierarchical Bayesian model enables inference regarding the location and date of introduction from indirect observation of the process such as species or disease occurrence records. We illustrate our approach using disease surveillance data collected to monitor white-nose syndrome, which is a fungal disease that threatens many North American species of bats. We use our model and surveillance data to estimate the location and date that the pathogen was introduced into the United States. Finally, we compare forecasts from our model to forecasts obtained from state-of-the-art regression-based statistical and machine learning methods. Our results show that the pathogen causing white-nose syndrome was most likely introduced into the United States 4 years prior to the first detection, but there is a moderate level of uncertainty in this estimate. The location of introduction could be up to 510 km east of the location of first discovery, but our results indicate that there is a relatively high probability the location of first detection could be the location of introduction.

stat.AP↗

Accounting for location uncertainty in distance sampling data

Ecologists use distance sampling to estimate the abundance of plants and animals while correcting for undetected individuals. By design, data collection is simplified by requiring only the distances from a transect to the detected individuals be recorded. Compared to traditional design-based methods that require restrictive assumption and limit the use of distance sampling data, model-based approaches enable broader applications such as spatial prediction, inferring species-habitat relationships, unbiased estimation from preferentially sampled transects, and integration into multi-type data models. Unfortunately, model-based approaches require the exact location of each detected individual in order to incorporate environmental and habitat characteristics as predictor variables. We modified model-based methods for distance sampling data by including a probability distribution that accounts for location uncertainty generated when only the distances are recorded. We tested and demonstrated our method using a simulation experiment and by modeling the habitat use of Dickcissels (Spiza americana) using distance sampling data collected from the Konza Prairie in Kansas, USA. Our results showed that ignoring location uncertainty can result in biased coefficient estimates and predictions. However, accounting for location uncertainty remedies the issue and results in reliable inference and prediction. Like other types of measurement error, hierarchical models can accommodate the data collection process thereby enabling reliable inference. Our approach is a significant advancement for the analysis of distance sampling data because it remedies the deleterious effects of location uncertainty and requires only distances be recorded. In turn, this enables historical distance sampling data sets to be compatible with modern data collection and modeling practices.

stat.AP↗

Animal Movement Models for Migratory Individuals and Groups

Animals often exhibit changes in their behavior during migration. Telemetry data provide a way to observe geographic position of animals over time, but not necessarily changes in the dynamics of the movement process. Continuous-time models allow for statistical predictions of the trajectory in the presence of measurement error and during periods when the telemetry device did not record the animal's position. However, continuous-time models capable of mimicking realistic trajectories with sufficient detail are computationally challenging to fit to large data sets and basic models lack realism in their ability to capture nonstationary dynamics. We present a unified class of animal movement models that are computationally efficient and provide a suite of approaches for accommodating nonstationarity in continuous trajectories due to migration and interactions among individuals. We show how to nest convolution models to incorporate interactions among migrating individuals to account for nonstationarity and provide inference about dynamic migratory networks. We demonstrate these approaches in two case studies involving migratory birds. Specifically, we used process convolution models with temporal deformation to account for heterogeneity in individual greater white-fronted goose migrations in Europe and Iceland and we used nested process convolutions to model dynamic migratory networks in sandhill cranes in North America. The approach we present accounts for various forms of temporal heterogeneity in animal movement and is not limited to migratory applications. Furthermore, our models rely on well-established principles for modeling dependent data and leverage modern approaches for modeling dynamic networks to help explain animal movement and social interaction.

stat.AP↗

The basis function approach for modeling autocorrelation in ecological data

Analyzing ecological data often requires modeling the autocorrelation created by spatial and temporal processes. Many of the statistical methods used to account for autocorrelation can be viewed as regression models that include basis functions. Understanding the concept of basis functions enables ecologists to modify commonly used ecological models to account for autocorrelation, which can improve inference and predictive accuracy. Understanding the properties of basis functions is essential for evaluating the fit of spatial or time-series models, detecting a hidden form of multicollinearity, and analyzing large data sets. We present important concepts and properties related to basis functions and illustrate several tools and techniques ecologists can use when modeling autocorrelation in ecological data.

stat.AP↗