SearcharxivSearch

arXiv subjects

Alan E. Gelfand

Publications and source records attributed to Alan E. Gelfand.

At least 19 recordsLinked to original sources

Joint Spatial and Temporal Generalized Dissimilarity Mixed Modeling (stGDMM) for Beta Diversity

Generalized dissimilarity models (GDMs) have emerged as a valuable tool for formal statistical analysis of biodiversity. In particular, beta diversity, measured using dissimilarity measures, e.g., the Bray-Curtis dissimilarity in our case, provides a statistical summary of the difference in species composition between sites. It also provides novel data for spatial and spatio-temporal modeling as it resides over the product space of one space-time pair and a second space-time pair. In earlier work we developed the spatial generalized dissimilarity mixed model (spGDMM) to remedy some of the stochastic issues concerned with the foundational GDM in the literature. Here, we extend that work to include dynamics. We find much richer modeling opportunities as we consider beta diversity with regard to change in time as well as space. We illustrate with a dataset from the Cape Floristic Region (CFR) in South Africa.

stat.ME

Double zero-inflated spatio-temporal modeling of daily precipitation under detection thresholds

Explaining precipitation behavior at daily scale is important for fine scale understanding of the mechanisms driving precipitation. However, this effort is challenging because of the frequent incidence of zeros. The challenge is amplified by the acknowledged incidence of two types of zeros -- absence of precipitation as a dry event and absence of measured precipitation due to detection limits. In this work, we propose a multilevel spatio-temporal model which allows us to distinguish and explain the two types of zeros, as well as to model positive precipitation above the detection limit. The methodology combines a point mass at zero with probability modeled through a probit regression, a Gamma regression for latent positive precipitation amounts, and an observation mechanism subject to threshold-induced censoring. To capture spatial dependencies, Gaussian processes are employed in each regression model. Working within a Bayesian framework, we can obtain a rich range of inference with exact uncertainty. In particular, we provide model-based inference tools to compare and quantify differences between the true precipitation process and its observed counterpart across relevant characteristics. We apply our model to the analysis of daily spring observations at 70 sites over 15 years from the Ebro River Basin in northeastern Spain. Our findings indicate that the threshold strongly affects the occurrence of observed precipitation, especially in humid regions. While its impact on total accumulated amounts is small, it can exert a relevant effect on upper quantiles.

stat.ME

Analyzing spatial point processes degraded by displacement and imperfect detection

Spatial point processes are a valuable tool for probabilistic modeling to explain location data. However, the data themselves are often observed imperfectly. In order to perform accurate inference, one must account for these imperfections, which we refer to as degradation. We consider two forms of degradation for spatial Poisson processes: thinning and displacement. First, we provide some theoretical results on model identifiability, showing that, under weak conditions, one can jointly learn the scale of the displacement, a parametric form of thinning, and a nonparametric intensity function. The ability to learn all of these components and the resulting improvements for inference compared to the conceptual non-degraded but misspecified model are shown empirically via simulation study. Finally, we apply this approach to North Atlantic right whale call data from Cape Cod Bay.

stat.ME

Prior Smoothing for Multivariate Disease Mapping Models

To date, we have seen the emergence of a large literature on multivariate disease mapping. That is, incidence of (or mortality from) multiple diseases is recorded at the scale of areal units where incidence (mortality) across the diseases is expected to manifest dependence. The modeling involves a hierarchical structure: a Poisson model for disease counts (conditioning on the rates) at the first stage, and a specification of a function of the rates using spatial random effects at the second stage. These random effects are specified as a prior and introduce spatial smoothing to the rate (or risk) estimates. What we see in the literature is the amount of smoothing induced under a given prior across areal units compared with the observed/empirical risks. Our contribution here extends previous research on smoothing in univariate areal data models. Specifically, for three different choices of multivariate prior, we investigate both within prior smoothing according to hyperparameters and across prior smoothing. Its benefit to the user is to illuminate the expected nature of departure from perfect fit associated with these priors since model performance is not a question of goodness of fit. We propose both theoretical and empirical metrics for our investigation and illustrate with both simulated and real data.

stat.ME

Modeling Animal Communication Using Multivariate Hawkes Processes with Additive Excitation and Multiplicative Inhibition

Animal acoustic communication often exhibits temporal dependence, with calls triggering or suppressing subsequent calls within and across call types, individuals, or species. While Hawkes processes provide a natural framework for modeling excitation, incorporating inhibition in multivariate settings can raise identifiability issues and complicate parameter interpretation. We propose a flexible class of multivariate Hawkes processes that combines additive excitation with multiplicative inhibition. This formulation preserves the branching process interpretation of excitation while reducing confounding between excitation and inhibition, and allows direct quantification of background and excitation contributions to the event rate. Bayesian inference is conducted via Markov chain Monte Carlo, and model adequacy is assessed using the random time change theorem. The proposed methodology is evaluated through simulation and applied to two acoustic communication datasets: group-living meerkats, for which we analyze three selected call types with distinct behavioral roles, and a two-species baleen whale dataset involving humpback and North Atlantic right whales. The meerkat analysis reveals significant within- and cross-type excitation with cross-type inhibition, whereas the whale data show evidence primarily of within-species excitation.

stat.AP

On prior smoothing with discrete spatial data in the context of disease mapping

Disease mapping attempts to explain observed health event counts across areal units, typically using Markov random field models. These models rely on spatial priors to account for variation in raw relative risk or rate estimates. Spatial priors introduce some degree of smoothing, wherein, for any particular unit, empirical risk or incidence estimates are either adjusted towards a suitable mean or incorporate neighbor-based smoothing. While model explanation may be the primary focus, the literature lacks a comparison of the amount of smoothing introduced by different spatial priors. Additionally, there has been no investigation into how varying the parameters of these priors influences the resulting smoothing. This study examines seven commonly used spatial priors through both simulations and real data analyses. Using areal maps of peninsular Spain and England, we analyze smoothing effects using two datasets with associated populations at risk. We propose empirical metrics to quantify the smoothing achieved by each model and theoretical metrics to calibrate the expected extent of smoothing as a function of model parameters. We employ areal maps in order to quantitatively characterize the extent of smoothing within and across the models as well as to link the theoretical metrics to the empirical metrics.

stat.ME

Joint space-time modelling for upper daily maximum and minimum temperature record-breaking

Record-breaking temperature events are now frequently in the news, proffered as evidence of climate change, and often bring significant economic and human impacts. Our previous work undertook the first substantial spatial modelling investigation of temperature record-breaking across years for any given day within the year, employing a dataset consisting of over sixty years of daily maximum temperatures across peninsular Spain. That dataset also supplies daily minimum temperatures (which, in fact, are now available through 2023). Here, the dataset is converted into a daily pair of binary events, indicators, for that day, of whether a yearly record was broken for the daily maximum temperature and/or for the daily minimum temperature. Joint modelling addresses several inference issues: (i) defining/modelling record-breaking with bivariate time series of yearly indicators, (ii) strength of relationship between record-breaking events, (iii) prediction of joint, conditional and marginal record-breaking, (iv) persistence in record-breaking across days, (v) spatial interpolation across peninsular Spain. We substantially expand our previous work to enable investigation of these issues. We observe strong correlation between both processes but a growing trend of climate change that is well differentiated between them both spatially and temporally as well as different strengths of persistence and spatial dependence.

stat.ME

Joint Spatiotemporal Modeling of Zooplankton and Whale Abundance in a Dynamic Marine Environment

North Atlantic right whales are an endangered species; their entire population numbers approximately 372 individuals, and they are subject to major anthropogenic threats. They feed on zooplankton species whose distribution shifts in a dynamic and warming oceanic environment. Because right whales in turn follow their shifting food resource, it is necessary to jointly study the distribution of whales and their prey. The innovative joint species distribution modeling (JSDM) contribution here is different from anything in the large JDSM literature, reflecting the processes and data we have to work with. Specifically, our JSDM supplies a geostatistical model for expected amount of zooplankton collected at a site. We require a point pattern model for the intensity of right whale abundance. The two process models are joined through a latent conditional-marginal specification. Further, each species has two data sources to inform their respective distributions and these sources require novel data fusion. What emerges is a complex multi-level model. Through simulation we demonstrate the ability of our joint specification to identify model unknowns and learn better about the species distributions than modeling them individually. We then apply our modeling to real data from Cape Cod Bay, Massachusetts in the U.S.

stat.AP

Markov modeling for a satellite tag data record of whale diving behavior

Cuvier's beaked whales (Ziphius cavirostris) are the deepest diving marine mammal, consistently diving to depths exceeding 1,000m for durations longer than an hour, making them difficult animals to study. They are important to study because they are sensitive to disturbances from naval sonar. Satellite-linked telemetry devices provide up to 14-day long records of dive behavior. However, the time series of depths is discretized to coarse bins due to bandwidth limitations. We analyze telemetry data from beaked whales that were exposed to moderate levels of sonar within controlled exposure experiments (CEEs) to study behavioral responses to sound exposure. We model the data as a hidden Markov model (HMM) over the time series of discrete depth bins, introducing partially observed movement types and recent diving activity covariates to model marginal non-stationarity. Movement types provide more flexible modeling for CEEs than partially observed dive stages, which are more commonly used in dive behavior HMMs. We estimate the proposed model within a hierarchical Bayesian framework, using HMM methods to compute marginalized likelihoods and posterior predictive distributions. We assess behavioral response by comparing observed post-exposure behavior to usual unexposed behavior via the posterior predictive distribution. The model quantifies patterns in baseline diving behavior and finds evidence that beaked whales deviate in response to sound. We find evidence that (i) beaked whales initially shorten the time they spend between deep dives, which may have physiological effects and (ii) subsequently avoid deep dives, which can result in lost foraging opportunities.

stat.AP

Analyzing whale calling through Hawkes process modeling

Sound is assumed to be the primary modality of communication among marine mammal species. Analyzing acoustic recordings helps to understand the function of the acoustic signals as well as the possible impact of anthropogenic noise on acoustic behavior. Motivated by a dataset from a network of hydrophones in Cape Cod Bay, Massachusetts, utilizing automatically detected calls in recordings, we study the communication process of the endangered North Atlantic right whale. For right whales an "up-call" is known as a contact call, and ensuing counter-calling between individuals is presumed to facilitate group cohesion. We present novel spatiotemporal excitement modeling consisting of a background process and a counter-call process. The background process intensity incorporates the influences of diel patterns and ambient noise on occurrence. The counter-call intensity captures potential excitement, that calling elicits calling behavior. Call incidence is found to be clustered in space and time; a call seems to excite more calls nearer to it in time and space. We find evidence that whales make more calls during twilight hours, respond to other whales nearby, and are likely to remain quiet in the presence of increased ambient noise.

stat.AP

Spatio-temporal modeling for record-breaking temperature events in Spain

Record-breaking temperature events are now very frequently in the news, viewed as evidence of climate change. With this as motivation, we undertake the first substantial spatial modeling investigation of temperature record-breaking across years for any given day within the year. We work with a dataset consisting of over sixty years (1960-2021) of daily maximum temperatures across peninsular Spain. Formal statistical analysis of record-breaking events is an area that has received attention primarily within the probability community, dominated by results for the stationary record-breaking setting with some additional work addressing trends. Such effort is inadequate for analyzing actual record-breaking data. Effective analysis requires rich modeling of the indicator events which define record-breaking sequences. Resulting from novel and detailed exploratory data analysis, we propose hierarchical conditional models for the indicator events. After suitable model selection, we discover explicit trend behavior, necessary autoregression, significance of distance to the coast, useful interactions, helpful spatial random effects, and very strong daily random effects. Illustratively, the model estimates that global warming trends have increased the number of records expected in the past decade almost two-fold, 1.93 (1.89,1.98), but also estimates highly differentiated climate warming rates in space and by season.

stat.ME

Assessing Marine Mammal Abundance: A Novel Data Fusion

Marine mammals are increasingly vulnerable to human disturbance and climate change. Their diving behavior leads to limited visual access during data collection, making studying the abundance and distribution of marine mammals challenging. In theory, using data from more than one observation modality should lead to better informed predictions of abundance and distribution. With focus on North Atlantic right whales, we consider the fusion of two data sources to inform about their abundance and distribution. The first source is aerial distance sampling which provides the spatial locations of whales detected in the region. The second source is passive acoustic monitoring (PAM), returning calls received at hydrophones placed on the ocean floor. Due to limited time on the surface and detection limitations arising from sampling effort, aerial distance sampling only provides a partial realization of locations. With PAM, we never observe numbers or locations of individuals. To address these challenges, we develop a novel thinned point pattern data fusion. Our approach leads to improved inference regarding abundance and distribution of North Atlantic right whales throughout Cape Cod Bay, Massachusetts in the US. We demonstrate performance gains of our approach compared to that from a single source through both simulation and real data.

stat.AP

Bayesian joint quantile autoregression

Quantile regression continues to increase in usage, providing a useful alternative to customary mean regression. Primary implementation takes the form of so-called multiple quantile regression, creating a separate regression for each quantile of interest. However, recently, advances have been made in joint quantile regression, supplying a quantile function which avoids crossing of the regression across quantiles. Here, we turn to quantile autoregression (QAR), offering a fully Bayesian version. We extend the initial quantile regression work of Koenker and Xiao (2006) in the spirit of Tokdar and Kadane (2012). We offer a directly interpretable parametric model specification for QAR. Further, we offer a p-th order QAR(p) version, a multivariate QAR(1) version, and a spatial QAR(1) version. We illustrate with simulation as well as a temperature dataset collected in Aragón, Spain.

stat.ME

Model-based tools for assessing space and time change in daily maximum temperature: an application to the Ebro basin in Spain

There is continuing interest in the investigation of change in temperature over space and time. We offer a set of tools to illuminate such change temporally, at desired temporal resolution, and spatially, according to region of interest, using data generated from suitable space-time models. These tools include predictive spatial probability surfaces and spatial extents for an event. Working with exceedance events around the center of the temperature distribution, the probability surfaces capture the spatial variation in the risk of an exceedance event, while the spatial extents capture the expected proportion of incidence of a given exceedance event for a region of interest. Importantly, the proposed tools can be used with the output from any suitable model fitted to any set of spatially referenced time series data. As an illustration, we employ a dataset from 1956 to 2015 collected at 18 stations over Aragón in Spain, and a collection of daily maximum temperature series obtained from posterior predictive simulation of a Bayesian hierarchical daily temperature model. The results for the summer period show that although there is an increasing risk in all the events used to quantify the effects of climate change, it is not spatially homogeneous, with the largest increase arising in the center of Ebro valley and Eastern Pyrenees area. The risk of an increase of the average temperature between 1966-1975 and 2006-2015 higher than $1^\circ$C is higher than 0.5 all over the region, and close to 1 in the previous areas. The extent of daily temperature higher than the reference mean has increased 3.5% per decade. The mean of the extent indicates that 95% of the area under study has suffered a positive increment of the average temperature, and almost 70% higher than $1^{\circ}$C.

stat.ME

Joint Multivariate and Functional Modeling for Plant Traits and Reflectances

The investigation of leaf-level traits in response to varying environmental conditions has immense importance for understanding plant ecology. Remote sensing technology enables measurement of the reflectance of plants to make inferences about underlying traits along environmental gradients. While much focus has been placed on understanding how reflectance and traits are related at the leaf-level, the challenge of modelling the dependence of this relationship along environmental gradients has limited this line of inquiry. Here, we take up the problem of jointly modeling traits and reflectance given environment. Our objective is to assess not only response to environmental regressors but also dependence between trait levels and the reflectance spectrum in the context of this regression. This leads to joint modeling of a response vector of traits with reflectance arising as a functional response over the wavelength spectrum. To conduct this investigation, we employ a dataset from a global biodiversity hotspot, the Greater Cape Floristic Region in South Africa.

stat.AP

Zero-inflated Beta distribution regression modeling

A frequent challenge encountered with ecological data is how to interpret, analyze, or model data having a high proportion of zeros. Much attention has been given to zero-inflated count data, whereas models for non-negative continuous data with an abundance of 0s are lacking. We consider zero-inflated data on the unit interval and provide modeling to capture two types of 0s in the context of the Beta regression model. We model 0s due to missing by chance through left censoring of a latent regression, and 0s due to unsuitability using an independent Bernoulli specification to create a point mass at 0. We first develop the model as a spatial regression in environmental features and then extend to introduce spatial random effects. We specify models hierarchically, employing latent variables, fit them within a Bayesian framework, and present new model comparison tools. Our motivating dataset consists of percent cover abundance of two plant species at a collection of sites in the Cape Floristic Region of South Africa. We find that environmental features enable learning about the incidence of both types of 0s as well as the positive percent covers. We also show that the spatial random effects model improves predictive performance. The proposed modeling enables ecologists, using environmental regressors, to extract a better understanding of the presence/absence of species in terms of absence due to unsuitability vs. missingness by chance, as well as abundance when present.

stat.ME

Spatial modeling of day-within-year temperature time series: an examination of daily maximum temperatures in Aragón, Spain

Acknowledging a considerable literature on modeling daily temperature data, we propose a multi-level spatio-temporal model which introduces several innovations in order to explain the daily maximum temperature in the summer period over 60 years in a region containing Aragón, Spain. The model operates over continuous space but adopts two discrete temporal scales, year and day within year. It captures temporal dependence through autoregression on days within year and also on years. Spatial dependence is captured through spatial process modeling of intercepts, slope coefficients, variances, and autocorrelations. The model is expressed in a form which separates fixed effects from random effects and also separates space, years, and days for each type of effect. Motivated by exploratory data analysis, fixed effects to capture the influence of elevation, seasonality and a linear trend are employed. Pure errors are introduced for years, for locations within years, and for locations at days within years. The performance of the model is checked using a leave-one-out cross-validation. Applications of the model are presented including prediction of the daily temperature series at unobserved or partially observed sites and inference to investigate climate change comparison.

stat.ME

Preferential Sampling for Bivariate Spatial Data

Preferential sampling provides a formal modeling specification to capture the effect of bias in a set of sampling locations on inference when a geostatistical model is used to explain observed responses at the sampled locations. In particular, it enables modification of spatial prediction adjusted for the bias. Its original presentation in the literature addressed assessment of the presence of such sampling bias while follow on work focused on regression specification to improve spatial interpolation under such bias. All of the work in the literature to date considers the case of a univariate response variable at each location, either continuous or modeled through a latent continuous variable. The contribution here is to extend the notion of preferential sampling to the case of bivariate response at each location. This exposes sampling scenarios where both responses are observed at a given location as well as scenarios where, for some locations, only one of the responses is recorded. That is, there may be different sampling bias for one response than for the other. It leads to assessing the impact of such bias on co-kriging. It also exposes the possibility that preferential sampling can bias inference regarding dependence between responses at a location. We develop the idea of bivariate preferential sampling through various model specifications and illustrate the effect of these specifications on prediction and dependence behavior. We do this both through simulation examples as well as with a forestry dataset that provides mean diameter at breast height (MDBH) and trees per hectare (TPH) as the point-referenced bivariate responses.

stat.ME