Searcharxiv⌕ Search

arXiv subjects

Jonathan R. Bradley

Publications and source records attributed to Jonathan R. Bradley.

At least 19 recordsLinked to original sources

Bayesian Analysis Using a Constrained Mixture of Normal-Inverse-Gamma Models

Gaussian mixtures of regressions are commonly implemented via a Gibbs sampler. This Markov chain Monte Carlo (MCMC) algorithm can be computationally burdensome because of the need to update discrete-valued latent component allocation parameters whose dimension increases as the sample size increases. In this article, we propose applying the method of composition to a Gaussian finite mixture model with a Normal-Inverse-Gamma (NIG) prior which allows one to write the posterior distribution as the product of conditional distributions. Namely, the conditional distribution of parameters given the data and mixture labels, times the marginal posterior of the mixture labels. The conditional distribution of parameters given the data and mixture labels, can be sampled from directly, instead of using MCMC. The expression of the marginal posterior of the mixture labels is known up to a proportionality constant and we adapt existing approaches in Bayesian selective inference to constrain the space of component labels to those arising from preliminary estimators, which alleviates a commonly encountered bottleneck. In simulation studies, we consider several settings and compare several versions of our constrained mixture of NIG models to two different MCMC-based strategies and demonstrate their use on natality data from the CDC.

stat.ME↗

A Bayesian Approach to Unit-level Dependent Multi-type Survey Data

The American Community Survey (ACS) Public Use Microdata Sample (PUMS) provides access to a wide range of unit-level survey data consisting of correlated Gaussian and binomial distributed survey responses along with associated survey weights. As such, we propose a Bayesian hierarchical framework for jointly modeling unit-level Gaussian and binomial survey data. The model introduces a shared area-level random effect to capture dependence across responses. Informative sampling is addressed using a pseudo-likelihood construction, and Polya-Gamma data augmentation provides an efficient conjugate Gibbs sampler, enabling scalable inference for large survey datasets. Through empirical simulations based on ACS PUMS data, we show that the joint model achieves notable reductions in mean squared error and improved interval scores compared to univariate and design-based estimators. Applying the method to the 2023 Illinois PUMS data, we find that the joint model yields small-area estimates similar to those from the univariate model and the Horvitz-Thompson estimator, but with smaller posterior variances. The computational cost associated with the joint model is also comparable to that of the univariate binomial model. Combined with the empirical simulation results, these findings demonstrate the practical advantages of the proposed approach.

stat.ME↗

Restricted Spatial Regression is Reasonable Statistical Practice: Clarifications, Interpretations, and New Developments

The spatial linear mixed model (SLMM) consists of fixed and spatial random effects that may be linearly dependent. Partially motivated as a means to address potential issues with confounding, the Restricted spatial regression (RSR) model restricts spatial random effects to be in the orthogonal column space of the covariates. Recent articles have shown that the misspecified Bayesian RSR generally performs worse than the SLMM when the data is generated from the SLMM. However, we show that the misspecified Bayesian RSR model's marginal posterior distribution is equivalent up to a reparameterization to that of the SLMM's marginal posterior distribution, under a certain prior assumption on the orthogonalized regression coefficients. This suggests that the RSR models are not sub-optimal as the subsequent Bayesian analysis can be interpreted as a type of SLMM Bayesian analysis. This equivalence relationship is developed further in the context of unmeasured confounders and nonlinearity, where we explore a semi-parametric property of the orthogonalized regression effects. Several results are provided to demonstrate new benefits of an RSR. In particular, we provide new results that show that the RSR can produce clear computational advantages via a direct sampler from the posterior distribution for all hyperparameters, fixed effects, and random effects. Additionally, a transfer learning approach offers a new interpretation to orthogonalized regression coefficients, which we show empirically can improve inference on dependent regression coefficients in the presence of spatial confounding. Simulations and an illustration using COVID-19 mortality data are provided.

stat.ME↗

Bayesian Inference for Spatial-Temporal Non-Gaussian Data Using Predictive Stacking

Analysing non-Gaussian spatial-temporal data requires introducing spatial as well as temporal dependence in generalised linear models through the link function of an exponential family distribution. Unlike in Gaussian likelihoods, inference is considerably encumbered by the inability to analytically integrate out the random effects and reduce the dimension of the parameter space. Iterative estimation algorithms struggle to converge due to the presence of weakly identified parameters. We devise Bayesian inference using predictive stacking that assimilates inference from analytically tractable conditional posterior distributions. We achieve this by expanding upon the Diaconis-Ylvisaker family of conjugate priors and exploiting generalised conjugate multivariate (GCM) distribution theory for exponential families, which enables exact sampling from analytically available posterior distributions conditional upon some process parameters. Subsequently, we assimilate inference over a range of values of these parameters using Bayesian predictive stacking. We evaluate inferential performance on simulated data, compare with full Bayesian inference using Markov chain Monte Carlo (MCMC) and apply our method to analyse spatially-temporally referenced avian count data from the North American Breeding Bird Survey database.

stat.ME↗

An Efficient Class of Bayesian Generalized Quadratic Nonlinear Dynamic Models with Application to Birth Rate Monitoring

Many real-world spatio-temporal processes exhibit nonlinear dynamics that can often be described through stochastic partial differential equations. These models are flexible and scientifically motivated, however, implementing them in a fully Bayesian framework can be computationally challenging. We are motivated by birth rate data, which has important implications for public health and are known to follow nonlinear dynamics. We propose a covariance calibration strategy that specifies the covariance matrix of a linear mixed effects model to be close in Frobenius norm to that of a Generalized Quadratic Nonlinearity (GQN) model. We refer to this as Frobenius norm matching. This allows us to model nonlinear dynamics using an easier to implement linear framework. The calibrated linear model is efficiently implemented using Exact Posterior Regression (EPR), a recently proposed Bayesian model that enables sampling of fixed and random effects directly from the posterior distribution. We provide simulation studies that compare to implementations using MCMC. Finally, we use this approach to analyze Florida county-level birth rate data from 1990-2023. Our results indicate that our non-linear spatio-temporal model outperforms linear dynamic spatio-temporal models for this data, and identifies covariate effects consistent with existing literature, all while avoiding the computational difficulties of MCMC.

stat.ME↗

Exact Bayesian Inference for Multivariate Spatial Data of Any Size with Application to Air Pollution Monitoring

Fine particulate matter and aerosol optical thickness are of interest to atmospheric scientists for understanding air quality and its various health/environmental impacts. The available data are extremely large, making uncertainty quantification in a fully Bayesian framework quite difficult, as traditional implementations do not scale reasonably to the size of the data. We specifically consider roughly 8 million observations obtained from NASA's Moderate Resolution Imaging Spectroradiometer (MODIS) instrument. To analyze data on this scale, we introduce Scalable Multivariate Exact Posterior Regression (SM-EPR) which combines the recently introduced data subset approach and Exact Posterior Regression (EPR). EPR is a new Bayesian hierarchical model where it is possible to sample independent replicates of fixed and random effects directly from the posterior without the use of Markov chain Monte Carlo (MCMC). We extend EPR to the multivariate spatial context, where the multiple variables may be distributed according to different distributions. The combination of the data subset approach with EPR allows one to perform exact Bayesian inference without MCMC for effectively any sample size. Additional motivation is provided via technical results illustrating favorable Kullback-Leibler and covariance properties. We demonstrate SM-EPR using a motivating big remote sensing data application and provide several simulations.

stat.ME↗

A Criterion for Aggregation Error for Multivariate Spatial Data

The criterion for aggregation error (CAGE) is an important metric that aims to measure errors that arise in multiscale (or multi-resolution) spatial data, referred to as the modifiable areal unit problem and the ecological fallacy. Specifically, CAGE is a measure of between scale variance of eigenvectors in a Karhunen-Loéve expansion (KLE), motivated by a theoretical result, referred to as the ``null-MAUP-theorem,'' that states that the MAUP/ecological fallacy are not present when this variance is zero. CAGE was originally developed for univariate spatial data, but its use has been applied to multivariate spatial data without the development of a null-MAUP-theorem in the multivariate spatial setting. To fill this gap, we provide theoretical justification for a multivariate CAGE (MVCAGE), which includes multiscale multivariate extensions of the KLE, Mercer's theorem, and the-null-MAUP theorem. Additionally, we provide technical results that demonstrate that the MVCAGE is preferable to spatial-only CAGE, and extend commonly used basis functions used to compute CAGE to the multivariate spatial setting. Empirical results are provided to demonstrate the use of MVCAGE for uncertainty quantification and regionalization.

stat.ME↗

Multiscale Multi-Type Spatial Bayesian Analysis for High-Dimensional Data with Application to Wildfires and Migration

Wildfires have significantly increased in the United States (U.S.), making certain areas harder to live in. This motivates us to jointly analyze active fires and population changes in the U.S. from July 2020 to June 2021. The available data are recorded on different scales (or spatial resolutions) and by different types of distributions (referred to as multi-type data). Moreover, wildfires are known to have feedback mechanism that creates signal-to-noise dependence. We analyze point-referenced remote sensing fire data from National Aeronautics and Space Administration (NASA) and county-level population change data provided by U.S. Census Bureau's Population Estimates Program (PEP). We develop a multiscale multi-type spatial Bayesian model that assumes the average number of fires is zero-inflated normal, the incidence of fire as Bernoulli, and the percentage population change as normally distributed. This high-dimensional dataset makes Markov chain Monte Carlo (MCMC) implementation infeasible. We bypass MCMC by extending a recently introduced computationally efficient Bayesian framework to directly sample from the exact posterior distribution, which includes a term to model signal-to-noise dependence. Such signal-to-noise dependence is known to be present in wildfire data, but is commonly not accounted for. A simulation study is used to highlight the computational performance of our method. In our analysis, we obtained predictions of wildfire probabilities, identified several useful covariates, and found that regions with many fires were associated with population change.

stat.ME↗

Markov Random Fields with Proximity Constraints for Spatial Data

The conditional autoregressive (CAR) model, simultaneous autoregressive (SAR) model, and its variants have become the predominant strategies for modeling regional or areal-referenced spatial data. The overwhelming wide-use of the CAR/SAR model motivates the need for new classes of models for areal-referenced data. Thus, we develop a novel class of Markov random fields based on truncating the full-conditional distribution. We define this truncation in two ways leading to versions of what we call the truncated autoregressive (TAR) model. First, we truncate the full conditional distribution so that a response at one location is close to the average of its neighbors. This strategy establishes relationships between TAR and CAR. Second, we truncate on the joint distribution of the data process in a similar way. This specification leads to connection between TAR and SAR model. Our Bayesian implementation does not use Markov chain Monte Carlo (MCMC) for Bayesian computation, and generates samples directly from the posterior distribution. Moreover, TAR does not have a range parameter that arises in the CAR/SAR models, which can be difficult to learn. We present the results of the proposed truncated autoregressive model on several simulated datasets and on a dataset of average property prices.

stat.ME↗

Bayesian Hierarchical Modeling for Bivariate Multiscale Spatial Data with Application to Blood Test Monitoring

In public health applications, spatial data collected are often recorded at different spatial scales and over different correlated variables. Spatial change of support is a key inferential problem in these applications and have become standard in univariate settings; however, it is less standard in multivariate settings. There are several existing multivariate spatial models that can be easily combined with multiscale spatial approach to analyze multivariate multiscale spatial data. In this paper, we propose three new models from such combinations for bivariate multiscale spatial data in a Bayesian context. In particular, we extend spatial random effects models, multivariate conditional autoregressive models, and ordered hierarchical models through a multiscale spatial approach. We run simulation studies for the three models and compare them in terms of prediction performance and computational efficiency. We motivate our models through an analysis of 2015 Texas annual average percentage receiving two blood tests from the Dartmouth Atlas Project.

stat.ME↗

Incorporating Subsampling into Bayesian Models for High-Dimensional Spatial Data

Additive spatial statistical models with weakly stationary process assumptions have become standard in spatial statistics. However, one disadvantage of such models is the computation time, which rapidly increases with the number of data points. The goal of this article is to apply an existing subsampling strategy to standard spatial additive models and to derive the spatial statistical properties. We call this strategy the ''spatial data subset model'' (SDSM) approach, which can be applied to big datasets in a computationally feasible way. Our approach has the advantage that one does not require any additional restrictive model assumptions. That is, computational gains increase as model assumptions are removed when using our model framework. This provides one solution to the computational bottlenecks that occur when applying methods such as Kriging to ''big data''. We provide several properties of this new spatial data subset model approach in terms of moments, sill, nugget, and range under several sampling designs. An advantage of our approach is that it subsamples without throwing away data, and can be implemented using datasets of any size that can be stored. We present the results of the spatial data subset model approach on simulated datasets, and on a large dataset consists of 150,000 observations of daytime land surface temperatures measured by the MODIS instrument onboard the Terra satellite.

stat.ME↗

Generating Independent Replicates Directly from the Posterior Distribution for a Class of Spatial Latent Gaussian Process Models

Markov chain Monte Carlo (MCMC) allows one to generate dependent replicates from a posterior distribution for effectively any Bayesian hierarchical model. However, MCMC can produce a significant computational burden. This motivates us to consider finding expressions of the posterior distribution that are computationally straightforward to obtain independent replicates from directly. We focus on a broad class of Bayesian latent Gaussian process (LGP) models that allow for spatially dependent data. First, we derive a new class of distributions we refer to as the generalized conjugate multivariate (GCM) distribution. The GCM distribution's theoretical development is similar to that of the CM distribution with two main differences; namely, (1) the GCM allows for latent Gaussian process assumptions, and (2) the GCM explicitly accounts for hyperparameters through marginalization. The development of GCM is needed to obtain independent replicates directly from the exact posterior distribution, which has an efficient projection/regression form. Hence, we refer to our method as Exact Posterior Regression (EPR). Illustrative examples are provided including simulation studies for weakly stationary spatial processes and spatial basis function expansions. An additional analysis of poverty incidence data from the U.S. Census Bureau's American Community Survey (ACS) using a conditional autoregressive model is presented.

stat.ME↗

Bayesian Hierarchical Models For Multi-type Survey Data Using Spatially Correlated Covariates Measured With Error

We introduce Bayesian hierarchical models for predicting high-dimensional tabular survey data which can be distributed from one or multiple classes of distributions (e.g., Gaussian, Poisson, Binomial, etc.). We adopt a Bayesian implementation of a Hierarchical Generalized Transformation (HGT) model to deal with the non-conjugacy of non-Gaussian data models when estimated using a Latent Gaussian Process (LGP) model. Survey data are usually prone to a high degree of sampling error, and we use covariates that are prone to measurement error as well as those free of any such error. A classical measurement error component is defined to deal with the sampling error in the covariates. The proposed models can be high-dimensional and we employ the notion of basis function expansions to provide an effective approach to dimension reduction. The HGT component lends flexibility to our model to incorporate multi-type response datasets under a unified latent process model framework. To demonstrate the applicability of our methodology, we provide the results from simulation studies and data applications arising from a dataset consisting of the U.S. Census Bureau's American Community Survey (ACS) 5-year period estimates of the total population count under the poverty threshold and the ACS 5-year period estimates of median housing costs at the county level across multiple states in the USA.

stat.ME↗

Interpolating Population Distributions using Public-use Data: An Application to Income Segregation using American Community Survey Data

Income segregation measures the extent to which households choose to live near other households with similar incomes. Sociologists theorize that income segregation can exacerbate the impacts of income inequality, and have developed indices to measure it at the metro area level, including the information theory index introduced in \citet{reardon2011income}, and the divergence index presented in \citet{roberto2015divergence}. To study their differences, we construct both indices using recent American Community Survey (ACS) estimates of features of the income distribution. Since the elimination of the decennial census long form, methods of computing these estimates must be updated to use ACS estimates and account for survey error. We propose a model-based method to interpolate estimates of features of the income distribution that accounts for this error. This method improves on previous approaches by allowing for the use of more types of estimates, and by providing uncertainty quantification. We apply this method to estimate U.S. census tract-level income distributions using ACS tabulations, and in turn use these to construct both income segregation indices. We find major differences between the two indices in the relative ranking of metro areas, as well as differences in how both indices correlate with the Gini index.

stat.ME↗

Constrained Bayesian Hierarchical Models for Gaussian Data: A Model Selection Criterion Approach

Consider the setting where there are B>1 candidate statistical models, and one is interested in model selection. Two common approaches to solve this problem are to select a single model or to combine the candidate models through model averaging. Instead, we select a subset of the combined parameter space associated with the models. Specifically, a model averaging perspective is used to increase the parameter space, and a model selection criterion is used to select a subset of this expanded parameter space. We account for the variability of the criterion by adapting Yekutieli (2012)'s method to Bayesian model averaging (BMA). Yekutieli (2012)'s method treats model selection as a truncation problem. We truncate the joint support of the data and the parameter space to only include small values of the covariance penalized error (CPE) criterion. The CPE is a general expression that contains several information criteria as special cases. Simulation results show that as long as the truncated set does not have near zero probability, we tend to obtain lower mean squared error than BMA. Additional theoretical results are provided that provide the foundation for these observations. We apply our approach to a dataset consisting of American Community Survey (ACS) period estimates to illustrate that this perspective can lead to improvements of a single model.

stat.ME↗

Spatio-Temporal Change of Support Modeling with R

Spatio-temporal change of support methods are designed for statistical analysis on spatial and temporal domains which can differ from those of the observed data. Previous work introduced a parsimonious class of Bayesian hierarchical spatio-temporal models, which we refer to as STCOS, for the case of Gaussian outcomes. Application of STCOS methodology from this literature requires a level of proficiency with spatio-temporal methods and statistical computing which may be a hurdle for potential users. The present work seeks to bridge this gap by guiding readers through STCOS computations. We focus on the R computing environment because of its popularity, free availability, and high quality contributed packages. The stcos package is introduced to facilitate computations for the STCOS model. A motivating application is the American Community Survey (ACS), an ongoing survey administered by the U.S. Census Bureau that measures key socioeconomic and demographic variables for various populations in the United States. The STCOS methodology offers a principled approach to compute model-based estimates and associated measures of uncertainty for ACS variables on customized geographies and/or time periods. We present a detailed case study with ACS data as a guide for change of support analysis in R, and as a foundation which can be customized to other applications.

stat.CO↗

Joint spatio-temporal analysis of multiple response types using the hierarchical generalized transformation model with application to coronavirus disease 2019 and social distancing

Social distancing can be described as an effort to maintain a physical distance between individuals and has become a necessary public health measure to combat cornoavirus disease 2019 (COVID-19). Social distancing is known to weaken incidences and deaths due to COVID-19, however, there are detrimental economic and psychological effects. This motivates us to analyze incidences (and deaths) of COVID-19 along with a measure of the health of the US economy (i.e., the adjusted closing price of the Dow Jones Industrial), and a measure of the public interest in COVID-19 through Google Trends data. The model we implement is developed to be easily adapted to a data scientist's preferred method for continuous data, which is done to aid future analyses of this important dataset. This dataset consists of multiple response types (e.g., continuous-valued, count-valued, binomial counts). Thus, we introduce a reasonable easy-to-implement all-purpose method that "converts" a statistical model for continuous responses (the preferred model) into a Bayesian model for multi-response data sets. To do this, we transform the data such that the continuous-valued transformed data can be reasonably modeled using the preferred model and the transformation itself is treated as unknown. The implementation of our approach involves two steps. The first step produces posterior replicates of the transformed data using a latent conjugate multivariate (LCM) model. The second step involves generating values from the posterior distribution implied by the preferred model. We refer to our model as the hierarchical generalized transformation (HGT) model. In a simulation, we demonstrate the flexibility of the HGT model by incorporating two different preferred models: Bayesian additive regression trees (BART) and the spatial mixed effects (spatio-temporal mixed effects) models.

stat.ME↗

Bayesian Inference for Big Spatial Data Using Non-stationary Spectral Simulation

It is increasingly understood that the assumption of stationarity is unrealistic for many spatial processes. In this article, we combine dimension expansion with a spectral method to model big non-stationary spatial fields in a computationally efficient manner. Specifically, we use Mejia and Rodriguez-Iturbe (1974)'s spectral simulation approach to simulate a spatial process with a covariogram at locations that have an expanded dimension. We introduce Bayesian hierarchical modelling to dimension expansion, which originally has only been modeled using a method of moments approach. In particular, we simulate from the posterior distribution using a collapsed Gibbs sampler. Our method is both full rank and non-stationary, and can be applied to big spatial data because it does not involve storing and inverting large covariance matrices. Additionally, we have fewer parameters than many other non-stationary spatial models. We demonstrate the wide applicability of our approach using a simulation study, and an application using ozone data obtained from the National Aeronautics and Space Administration (NASA).

stat.ME↗