SearcharxivSearch

arXiv subjects

Grace S. Chiu

Publications and source records attributed to Grace S. Chiu.

10 recordsLinked to original sources

Bayesian Gaussian Mixture Modeling for Symmetric Matrix Variate Data

Statistical inference on individual activity networks has been a historically difficult task due to the lack of available data at the appropriate granularity and the complexity of modeling individual mobility patterns. The recent availability of GPS data from individual devices, combined with highly detailed demographic information, suggests that one of these challenges can now be addressed. We introduce a new model which we call the Symmetric Matrix-Variate Normal Mixture Model (STRUCTURED) to estimate how demographic traits influence changes in human activity networks, using sociomatrices that capture the probabilistic spatial overlap between individuals over time. We exploit the commutativity constraint inherent in the symmetric matrix-variate normal distribution to parameterize the column precision matrix as a polynomial of the row precision matrix, reducing the effective parameter space by an order of magnitude. We develop two variants of STRUCTURED: STRUCTURED-FP, which estimates the full polynomial, and STRUCTURED-RJ, which uses reversible-jump MCMC to select a reduced-order parameterization. Simulation studies demonstrate that STRUCTURED-RJ outperforms existing methods in sparse-data regimes, whereas STRUCTURED-FP is preferred when sample sizes are large. We apply the model to GPS-derived sociomatrices of 293 individuals in King County, WA, finding that local crime environments and youth employment density are the dominant demographic factors explaining variation in weekly activity overlap patterns.

stat.ME

Big shells, bigger data: cohort analysis of Chesapeake Bay Crassostrea virginica reefs

Oysters in Virginia Chesapeake Bay oyster reefs are "age-truncated", possibly due to a combination of historical overfishing, disease epizootics, environmental degradation, and climate change. Research has suggested that oysters exhibit resilience to environmental stressors; however, that evidence is based on the current limited understanding of oyster lifespan. Until this paper, the Virginia Oyster Stock Assessment and Replenishment Archive (VOSARA), a spatially and temporally expansive dataset (222 reefs across 2003-2023) of shell lengths (SL, mm), had yet to be examined comprehensively in the context of resilience. We develop a novel method using Gaussian mixture modeling (GMM) to identify the age groups in each reef using yearly SL data and then link those age groups over time to identify cohorts and estimate their lifespan. Sixty-four reefs (29%) are deemed to have sufficient data (at least 300 oysters sampled for a minimum of 8 consecutive years) for this analysis. We fit univariate GMMs for each year ($t$) and reef ($r$) for each of the seven river strata ($R$) to estimate 1) the mean and standard deviation of SL for each $a_{Rrt}$th age group, and 2) the mixture percentage of each $a_{Rrt}$th age group. We link age groups across time to infer age cohorts by developing a mechanistic algorithm that prevents the shrinking of shell length when an $a_{Rrt}$th group becomes an ($a_{R,r,t+1}$)th group. Our method shows promise in identifying oyster cohorts and estimating lifespan solely using SL data. Our results show signals of resiliency in almost all river systems: oyster cohorts live longer and grow larger in the mid-to-late 2010s compared to the early 2000s.

stat.AP

Model-based calibration of gear-specific fish abundance survey data as a change-of-support problem

For commercial and recreational fisheries of a wide-ranging species to be sustainable, abundance studies from neighboring regions should be unified. For the first time in the USA, a single research project to estimate the abundance of the Greater Amberjack {Seriola dumerili) is being undertaken at the continental scale. A major methodological challenge lies in 1) the difference in fish detection gears deployed by regional survey teams that produce gear-specific relative abundance indices, and 2) the unknown relationship between actual abundance and these indices. In this paper, we develop a conversion tool that is operationalized from a Bayesian hierarchical model in an inferential context akin to the change-of-support problem often encountered in large-scale spatial studies; though, the context here is to reconcile abundance data observed at various gear-specific scales. To this end, we consider a small calibration experiment in which 2 to 4 different underwater video camera types were simultaneously deployed on each of 21 boat trips. Alongside the suite of deployed cameras was also an acoustic echosounder that recorded fish signals along surrounding transects. Our modeling framework is used to derive calibration formulae for translating camera-specific relative indices to the actual abundance scale in surveys that deploy a single camera. Cross-validation is conducted using mark-recapture abundance estimates (only available for 10 trips, all observed at a single habitat type) and through a separate simulation study. We also briefly discuss the case when surveys pair one camera with the echosounder.

stat.ME

Modeling Human Spatial Mobility Patterns with the Lévy Flight Cluster Model

Despite the extensive collection of individual mobility data over the past decade, fueled by the widespread use of GPS-enabled personal devices, the existing statistical literature on estimating human spatial mobility patterns from temporally irregular location data remains limited. In this paper, we introduce the Lévy Flight Cluster Model (LFCM), a hierarchical Bayesian mixture model designed to analyze an individual's activity distribution. The LFCM can be utilized to determine probabilistic overlaps between individuals' activity patterns and serves as an anonymization tool to generate synthetic location data. We present our methodology using real-world human location data, demonstrating its ability to accurately capture the key characteristics of human movement.

stat.ME

A New Perspective to Fish Trajectory Imputation: A Methodology for Spatiotemporal Modeling of Acoustically Tagged Fish Data

The focus of this paper is a key component of a methodology for understanding, interpolating, and predicting fish movement patterns based on spatiotemporal data recorded by spatially static acoustic receivers. Unlike GPS trackers which emit satellite signals from the animal's location, acoustic receivers are akin to stationary motion sensors that record movements within their detection range. Thus, for periods of time, fish may be far from the receivers, resulting in the absence of observations. The lack of information on the fish's location for extended time periods poses challenges to the understanding of fish movement patterns, and hence, the identification of proper statistical inference frameworks for modeling the trajectories. As the initial step in our methodology, in this paper, we devise and implement a simulation-based imputation strategy that relies on both Markov chain and random-walk principles to enhance our dataset over time. This methodology will be generalizable and applicable to all fish species with similar migration patterns or data with similar structures due to the use of static acoustic receivers.

stat.CO

Latent Causal Socioeconomic Health Index

This research develops a model-based LAtent Causal Socioeconomic Health (LACSH) index at the national level. Motivated by the need for a holistic national well-being index, we build upon the latent health factor index (LHFI) approach that has been used to assess the unobservable ecological/ecosystem health. LHFI integratively models the relationship between metrics, latent health, and covariates that drive the notion of health. In this paper, the LHFI structure is integrated with spatial modeling and statistical causal modeling. Our efforts are focused on developing the integrated framework to facilitate the understanding of how an observational continuous variable might have causally affected a latent trait that exhibits spatial correlation. A novel visualization technique to evaluate covariate balance is also introduced for the case of a continuous policy (treatment) variable. Our resulting LACSH framework and visualization tool are illustrated through two global case studies on national socioeconomic health (latent trait), each with various metrics and covariates pertaining to different aspects of societal health, and the treatment variable being mandatory maternity leave days and government expenditure on healthcare, respectively. We validate our model by two simulation studies. All approaches are structured in a Bayesian hierarchical framework and results are obtained by Markov chain Monte Carlo techniques.

stat.ME

Spatiotemporal Modeling of Nursery Habitat Using Bayesian Inference: Environmental Drivers of Juvenile Blue Crab Abundance

Nursery grounds are favorable for growth and survival of juvenile fish and crustaceans through abundant food resources and refugia, and enhance secondary production of populations. While small-scale studies remain important tools to assess nursery value of habitats, targeted applications that unify survey data over large spatiotemporal scales are vital to generalize inference of nursery function, identify highly productive regions, and inform management strategies. Using 21 years of GIS and spatiotemporally indexed field survey data on potential nursery habitats, we constructed five Bayesian models with varying spatiotemporal dependence structures to infer nursery habitat value for juveniles of the blue crab C. sapidus within three tributaries in lower Chesapeake Bay. Out-of-sample predictions of juvenile counts from a fully nonseparable spatiotemporal model outperformed predictions from simpler models. Salt marsh surface area, turbidity, and their interaction showed the strongest associations (and positively) with abundance. Relative seagrass area, previously emphasized as the most valuable nursery in small spatial-scale studies, was not associated with abundance. Hence, we argue that salt marshes should be considered a key nursery habitat for blue crabs, even amidst extensive seagrass beds. Moreover, identification of nurseries should be based on investigations at broad spatiotemporal scales incorporating multiple potential nursery habitats, and on rigorously addressing spatiotemporal dependence.

stat.AP

Modeling National Latent Socioeconomic Health and Examination of Policy Effects via Causal Inference

This research develops a socioeconomic health index for nations through a model-based approach which incorporates spatial dependence and examines the impact of a policy through a causal modeling framework. As the gross domestic product (GDP) has been regarded as a dated measure and tool for benchmarking a nation's economic performance, there has been a growing consensus for an alternative measure---such as a composite `wellbeing' index---to holistically capture a country's socioeconomic health performance. Many conventional ways of constructing wellbeing/health indices involve combining different observable metrics, such as life expectancy and education level, to form an index. However, health is inherently latent with metrics actually being observable indicators of health. In contrast to the GDP or other conventional health indices, our approach provides a holistic quantification of the overall `health' of a nation. We build upon the latent health factor index (LHFI) approach that has been used to assess the unobservable ecological/ecosystem health. This framework integratively models the relationship between metrics, the latent health, and the covariates that drive the notion of health. In this paper, the LHFI structure is integrated with spatial modeling and statistical causal modeling, so as to evaluate the impact of a policy variable (mandatory maternity leave days) on a nation's socioeconomic health, while formally accounting for spatial dependency among the nations. We apply our model to countries around the world using data on various metrics and potential covariates pertaining to different aspects of societal health. The approach is structured in a Bayesian hierarchical framework and results are obtained by Markov chain Monte Carlo techniques.

stat.AP

A Statistical Social Network Model for Consumption Data in Food Webs

We adapt existing statistical modeling techniques for social networks to study consumption data observed in trophic food webs. These data describe the feeding volume (non-negative) among organisms grouped into nodes, called trophic species, that form the food web. Model complexity arises due to the extensive amount of zeros in the data, as each node in the web is predator/prey to only a small number of other trophic species. Many of the zeros are regarded as structural (non-random) in the context of feeding behavior. The presence of basal prey and top predator nodes (those who never consume and those who are never consumed, with probability 1) creates additional complexity to the statistical modeling. We develop a special statistical social network model to account for such network features. The model is applied to two empirical food webs; focus is on the web for which the population size of seals is of concern to various commercial fisheries.

stat.ME

Assessing the Health of Richibucto Estuary with the Latent Health Factor Index

The ability to quantitatively assess the health of an ecosystem is often of great interest to those tasked with monitoring and conserving ecosystems. For decades, research in this area has relied upon multimetric indices of various forms. Although indices may be numbers, many are constructed based on procedures that are highly qualitative in nature, thus limiting the quantitative rigour of the practical interpretations made from these indices. The statistical modelling approach to construct the latent health factor index (LHFI) was recently developed to express ecological data, collected to construct conventional multimetric health indices, in a rigorous quantitative model that integrates qualitative features of ecosystem health and preconceived ecological relationships among such features. This hierarchical modelling approach allows (a) statistical inference of health for observed sites and (b) prediction of health for unobserved sites, all accompanied by formal uncertainty statements. Thus far, the LHFI approach has been demonstrated and validated on freshwater ecosystems. The goal of this paper is to adapt this approach to modelling estuarine ecosystem health, particularly that of the previously unassessed system in Richibucto in New Brunswick, Canada. Field data correspond to biotic health metrics that constitute the AZTI marine biotic index (AMBI) and abiotic predictors preconceived to influence biota. We also briefly discuss related LHFI research involving additional metrics that form the infaunal trophic index (ITI). Our paper is the first to construct a scientifically sensible model to rigorously identify the collective explanatory capacity of salinity, distance downstream, channel depth, and silt-clay content --- all regarded a priori as qualitatively important abiotic drivers --- towards site health in the Richibucto ecosystem.

stat.AP