SearcharxivSearch

arXiv subjects

Scott H. Holan

Publications and source records attributed to Scott H. Holan.

At least 19 recordsLinked to original sources

Nonprobability Samples for Small Area Estimation: A Review and Comparative Simulation Study

Nonprobability samples (NPS) are attractive because they are less costly to collect, can provide substantially larger sample sizes, and may reach populations that traditional probability surveys do not. As response rates for traditional surveys fall, interest in NPS has grown rapidly within the field of survey statistics. These methods are especially relevant for small area estimation (SAE), where there is ever-present demand for estimates at fine geographic scales and detailed demographic domains. Despite rapid methodological development, there remains limited understanding of which approaches perform best under different conditions. In this paper, we review recent developments in NPS methodology, including the concept of data defect correlation (DDC) as a measure of data quality and as a tool for categorizing the various NPS methods. We then present a comprehensive simulation study that evaluates a range of NPS approaches under varying levels of DDC and extend several existing methods to the SAE setting.

stat.ME

Bayesian Modeling of Gibbs Point Processes via Basis Function Expansions

We present a hierarchical Bayesian framework for non-homogeneous pairwise interaction Gibbs point process models, where the global and local effect functions are modeled via basis function expansions. We further propose a testing procedure in order to assess complete spatial randomness. The proposed methodology is exemplified through two real benchmark data examples involving water striders and forest fires.

stat.ME

Scalable Joint Modeling of Dependent Multi-Type Survey Data for Small Area Estimation

We develop a Bayesian area-level small area estimation framework that jointly models binomial and Gaussian survey responses through shared spatial random effects. This work is motivated by the American Community Survey (ACS), which provides useful information that contributes to federal funding and policy making decisions, and often yields direct estimates with large standard errors in small domains. The proposed Multi-type model borrows strength across outcomes and spatial neighbors to improve the precision of the associated estimates. For the binomial component, Polya-Gamma data augmentation yields a conditionally Gaussian representation, while spatial basis functions provide dimension reduction for high-dimensional spatial data. Together, these features lead to closed-form conditional posteriors and, thus, an efficient Gibbs sampler. Through empirical simulations, we show that the proposed joint model improves estimation precision relative to independent Univariate models. Applying the method to ACS median income and poverty rate data, we find that the proposed Multi-type model yields similar point estimates but smaller posterior variances than the corresponding Univariate models.

stat.ME

A Bayesian Approach to Unit-level Dependent Multi-type Survey Data

The American Community Survey (ACS) Public Use Microdata Sample (PUMS) provides access to a wide range of unit-level survey data consisting of correlated Gaussian and binomial distributed survey responses along with associated survey weights. As such, we propose a Bayesian hierarchical framework for jointly modeling unit-level Gaussian and binomial survey data. The model introduces a shared area-level random effect to capture dependence across responses. Informative sampling is addressed using a pseudo-likelihood construction, and Polya-Gamma data augmentation provides an efficient conjugate Gibbs sampler, enabling scalable inference for large survey datasets. Through empirical simulations based on ACS PUMS data, we show that the joint model achieves notable reductions in mean squared error and improved interval scores compared to univariate and design-based estimators. Applying the method to the 2023 Illinois PUMS data, we find that the joint model yields small-area estimates similar to those from the univariate model and the Horvitz-Thompson estimator, but with smaller posterior variances. The computational cost associated with the joint model is also comparable to that of the univariate binomial model. Combined with the empirical simulation results, these findings demonstrate the practical advantages of the proposed approach.

stat.ME

Echo State Networks for Spatio-Temporal Area-Level Data

Spatio-temporal area-level datasets play a critical role in official statistics, providing valuable insights for policy-making and regional planning. Accurate modeling and forecasting of these datasets can be extremely useful for policymakers to develop informed strategies for future planning. Echo State Networks (ESNs) are efficient methods for capturing nonlinear temporal dynamics and generating forecasts. However, ESNs lack a direct mechanism to account for the neighborhood structure inherent in area-level data. Ignoring these spatial relationships can significantly compromise the accuracy and utility of forecasts. In this paper, we incorporate approximate graph spectral filters at the input stage of the ESN, thereby improving forecast accuracy while preserving the model's computational efficiency during training. We demonstrate the effectiveness of our approach using Eurostat's tourism occupancy dataset and show how it can support more informed decision-making in policy and planning contexts.

cs.LG

An Anytime Valid Test for Complete Spatial Randomness

A relevant question when analyzing spatial point patterns is that of spatial randomness. More specifically, before any model can be fit to a point pattern a first step is to test the data for departures from complete spatial randomness (CSR). Traditional techniques employ distance or quadrat counts based methods to test for CSR based on batched data. In this paper, we consider the practical scenario of testing for CSR when the data are available sequentially (i.e., online). We present a sequential testing methodology called as {\em PRe-process} that is based on e-values and is a fast, efficient and nonparametric method. Simulation experiments with the truth departing from CSR in two different scenarios show that the method is effective in capturing inhomogeneity over time. Two real data illustrations considering lung cancer cases in the Chorley-Ribble area, England from 1974 - 1983 and locations of earthquakes in the state of Oklahoma, USA from 2000 - 2011 demonstrate the utility of the PRe-process in sequential testing of CSR.

stat.OT

Variational Autoencoded Multivariate Spatial Fay-Herriot Models

Small area estimation models are essential for estimating population characteristics in regions with limited sample sizes, thereby supporting policy decisions, demographic studies, and resource allocation, among other use cases. The spatial Fay-Herriot model is one such approach that incorporates spatial dependence to improve estimation by borrowing strength from neighboring regions. However, this approach often requires substantial computational resources, limiting its scalability for high-dimensional datasets, especially when considering multiple (multivariate) responses. This paper proposes two methods that integrate the multivariate spatial Fay-Herriot model with spatial random effects, learned through variational autoencoders, to efficiently leverage spatial structure. Importantly, after training the variational autoencoder to represent spatial dependence for a given set of geographies, it may be used again in future modeling efforts, without the need for retraining. Additionally, the use of the variational autoencoder to represent spatial dependence results in extreme improvements in computational efficiency, even for massive datasets. We demonstrate the effectiveness of our approach using 5-year period estimates from the American Community Survey over all census tracts in California.

stat.ML

Bayesian Unit-level Modeling of Categorical Survey Data with a Longitudinal Design

Categorical response data are ubiquitous in complex survey applications, yet few methods model the dependence across different outcome categories when the response is ordinal. Likewise, few methods exist for the common combination of a longitudinal design and categorical data. By modeling individual survey responses at the unit-level, it is possible to capture both ordering information in ordinal responses and any longitudinal correlation. However, accounting for a complex survey design becomes more challenging in the unit-level setting. We propose a Bayesian hierarchical, unit-level, model-based approach for categorical data that is able to capture ordering among response categories, can incorporate longitudinal dependence, and accounts for the survey design. To handle computational scalability, we develop efficient Gibbs samplers with appropriate data augmentation as well as variational Bayes algorithms. Using public-use microdata from the Household Pulse Survey, we provide an analysis of an ordinal response that asks about the frequency of anxiety symptoms at the beginning of the COVID-19 pandemic. We compare both design-based and model-based estimators and demonstrate superior performance for the proposed approaches.

stat.ME

A Criterion for Aggregation Error for Multivariate Spatial Data

The criterion for aggregation error (CAGE) is an important metric that aims to measure errors that arise in multiscale (or multi-resolution) spatial data, referred to as the modifiable areal unit problem and the ecological fallacy. Specifically, CAGE is a measure of between scale variance of eigenvectors in a Karhunen-Loéve expansion (KLE), motivated by a theoretical result, referred to as the ``null-MAUP-theorem,'' that states that the MAUP/ecological fallacy are not present when this variance is zero. CAGE was originally developed for univariate spatial data, but its use has been applied to multivariate spatial data without the development of a null-MAUP-theorem in the multivariate spatial setting. To fill this gap, we provide theoretical justification for a multivariate CAGE (MVCAGE), which includes multiscale multivariate extensions of the KLE, Mercer's theorem, and the-null-MAUP theorem. Additionally, we provide technical results that demonstrate that the MVCAGE is preferable to spatial-only CAGE, and extend commonly used basis functions used to compute CAGE to the multivariate spatial setting. Empirical results are provided to demonstrate the use of MVCAGE for uncertainty quantification and regionalization.

stat.ME

Topic Modeling for Free-Response Text Data from a Complex Survey

Topic Modeling is a popular statistical tool commonly used on textual data to identify the hidden thematic structure in a document collection based on the distribution of words. Additionally, it can be used to cluster the documents, with clusters representing distinct topics. The Mixture of Unigrams (MoU) is a standard topic model for clustering document-term data and can be particularly useful for analyzing open-ended survey responses to extract meaningful information from the underlying topics. However, with complex survey designs, where data is often collected on individual (document) characteristics, it is essential to account for the sample design in order to avoid biased estimates. To address this issue, we propose the MoU model under informative sampling using a pseudolikelihood to account for the sample design in the model by incorporating survey weights. We evaluate the effectiveness of this approach through a simulation study and illustrate its application using two datasets from the American National Election Studies (ANES). We compare our pseudolikelihood-based MoU model to the traditional MoU and assess its effectiveness in extracting meaningful topics from survey data. Additionally, we introduce a hierarchical Mixture of Unigrams (hMoU) accounting for informative sampling where topic proportions are defined as functions of document-level fixed and random effects. We demonstrate the effectiveness of the proposed model through an application to ANES data comparing topic proportions across respondent-level factors such as gender, race, age group, and state.

stat.AP

Incorporating Asymmetric Loss for Real Estate Prediction with Area-level Spatial Data

We investigate two asymmetric loss functions, namely LINEX loss and power divergence loss for optimal spatial prediction with area-level data. With our motivation arising from the real estate industry, namely in real estate valuation, we use the Zillow Home Value Index (ZHVI) for county-level values to show the change in prediction when the loss is different (asymmetric) from a traditional squared error loss (symmetric) function. Additionally, we discuss the importance of choosing the asymmetry parameter, and propose a solution to this choice for a general asymmetric loss function. Since the focus is on area-level data predictions, we propose the methodology in the context of conditionally autoregressive (CAR) models. We conclude that choice of the loss functions for spatial area-level predictions can play a crucial role, and is heavily driven by the choice of parameters in the respective loss.

stat.AP

Bayesian Methods to Improve The Accuracy of Differentially Private Measurements of Constrained Parameters

Formal disclosure avoidance techniques are necessary to ensure that published data can not be used to identify information about individuals. The addition of statistical noise to unpublished data can be implemented to achieve differential privacy, which provides a formal mathematical privacy guarantee. However, the infusion of noise results in data releases which are less precise than if no noise had been added, and can lead to some of the individual data points being nonsensical. Examples of this are estimates of population counts which are negative, or estimates of the ratio of counts which violate known constraints. A straightforward way to guarantee that published estimates satisfy these known constraints is to specify a statistical model and incorporate a prior on census counts and ratios which properly constrains the parameter space. We utilize rejection sampling methods for drawing samples from the posterior distribution and we show that this implementation produces estimates of population counts and ratios which maintain formal privacy, are more precise than the original unconstrained noisy measurements, and are guaranteed to satisfy prior constraints.

stat.ME

The Link Between Health Insurance Coverage and Citizenship Among Immigrants: Bayesian Unit-Level Regression Modeling of Categorical Survey Data Observed with Measurement Error

Social scientists are interested in studying the impact that citizenship status has on health insurance coverage among immigrants in the United States. This can be done using data from the Survey of Income and Program Participation (SIPP); however, two primary challenges emerge. First, statistical models must account for the survey design in some fashion to reduce the risk of bias due to informative sampling. Second, it has been observed that survey respondents misreport citizenship status at nontrivial rates. This too can induce bias within a statistical model. Thus, we propose the use of a weighted pseudo-likelihood mixture of categorical distributions, where the mixture component is determined by the latent true response variable, in order to model the misreported data. We illustrate through an empirical simulation study that this approach can mitigate the two sources of bias attributable to the sample design and misreporting. Importantly, our misreporting model can be further used as a component in a deeper hierarchical model. With this in mind, we conduct an analysis of the relationship between health insurance coverage and citizenship status using data from the SIPP.

stat.ME

A Socio-Demographic Latent Space Approach to Spatial Data When Geography is Important but Not All-Important

Many models for spatial and spatio-temporal data assume that "near things are more related than distant things," which is known as the first law of geography. While geography may be important, it may not be all-important, for at least two reasons. First, technology helps bridge distance, so that regions separated by large distances may be more similar than would be expected based on geographical distance. Second, geographical, political, and social divisions can make neighboring regions dissimilar. We develop a flexible Bayesian approach for learning from spatial data which units are close in an unobserved socio-demographic space and hence which units are similar. As a by-product, the Bayesian approach helps quantify the relative importance of socio-demographic space relative to geographical space. To demonstrate the proposed approach, we present simulations along with an application to county-level data on median household income in the U.S. state of Florida.

stat.ME

Bayesian Unit-level Models for Longitudinal Survey Data under Informative Sampling: An Analysis of Expected Job Loss Using the Household Pulse Survey

The Household Pulse Survey (HPS), recently released by the U.S. Census Bureau, gathers timely information about the societal and economic impacts of coronavirus. The first phase of the survey was quickly launched one month after the beginning of the coronavirus pandemic and ran for 12 weeks. To track the immediate impact of the pandemic, individual respondents during this phase were re-sampled for up to three consecutive weeks. Motivated by expected job loss during the pandemic, using public-use microdata, this work proposes unit-level, model-based estimators that incorporate longitudinal dependence at both the response and domain level. In particular, using a pseudo-likelihood, we consider a Bayesian hierarchical unit-level, model-based approach for both Gaussian and binary response data under informative sampling. To facilitate construction of these model-based estimates, we develop an efficient Gibbs sampler. An empirical simulation study is conducted to compare the proposed approach to models that do not account for unit-level longitudinal correlation. Finally, using public-use HPS micro-data, we provide an analysis of "expected job loss" that compares both design-based and model-based estimators and demonstrates superior performance for the proposed model-based approaches.

stat.ME

Bayesian Hierarchical Models For Multi-type Survey Data Using Spatially Correlated Covariates Measured With Error

We introduce Bayesian hierarchical models for predicting high-dimensional tabular survey data which can be distributed from one or multiple classes of distributions (e.g., Gaussian, Poisson, Binomial, etc.). We adopt a Bayesian implementation of a Hierarchical Generalized Transformation (HGT) model to deal with the non-conjugacy of non-Gaussian data models when estimated using a Latent Gaussian Process (LGP) model. Survey data are usually prone to a high degree of sampling error, and we use covariates that are prone to measurement error as well as those free of any such error. A classical measurement error component is defined to deal with the sampling error in the covariates. The proposed models can be high-dimensional and we employ the notion of basis function expansions to provide an effective approach to dimension reduction. The HGT component lends flexibility to our model to incorporate multi-type response datasets under a unified latent process model framework. To demonstrate the applicability of our methodology, we provide the results from simulation studies and data applications arising from a dataset consisting of the U.S. Census Bureau's American Community Survey (ACS) 5-year period estimates of the total population count under the poverty threshold and the ACS 5-year period estimates of median housing costs at the county level across multiple states in the USA.

stat.ME

Conjugate Modeling Approaches for Small Area Estimation with Heteroscedastic Structure

Small area estimation has become an important tool in official statistics, used to construct estimates of population quantities for domains with small sample sizes. Typical area-level models function as a type of heteroscedastic regression, where the variance for each domain is assumed to be known and plugged in following a design-based estimate. Recent work has considered hierarchical models for the variance, where the design-based estimates are used as an additional data point to model the latent true variance in each domain. These hierarchical models may incorporate covariate information, but can be difficult to sample from in high-dimensional settings. Utilizing recent distribution theory, we explore a class of Bayesian hierarchical models for small area estimation that smooth both the design-based estimate of the mean and the variance. In addition, we develop a class of unit-level models for heteroscedastic Gaussian response data. Importantly, we incorporate both covariate information as well as spatial dependence, while retaining a conjugate model structure that allows for efficient sampling. We illustrate our methodology through an empirical simulation study as well as an application using data from the American Community Survey.

stat.ME

Bayesian Circular Lattice Filters for Computationally Efficient Estimation of Multivariate Time-Varying Autoregressive Models

Nonstationary time series data exist in various scientific disciplines, including environmental science, biology, signal processing, econometrics, among others. Many Bayesian models have been developed to handle nonstationary time series. The time-varying vector autoregressive (TV-VAR) model is a well-established model for multivariate nonstationary time series. Nevertheless, in most cases, the large number of parameters presented by the model results in a high computational burden, ultimately limiting its usage. This paper proposes a computationally efficient multivariate Bayesian Circular Lattice Filter to extend the usage of the TV-VAR model to a broader class of high-dimensional problems. Our fully Bayesian framework allows both the autoregressive (AR) coefficients and innovation covariance to vary over time. Our estimation method is based on the Bayesian lattice filter (BLF), which is extremely computationally efficient and stable in univariate cases. To illustrate the effectiveness of our approach, we conduct a comprehensive comparison with other competing methods through simulation studies and find that, in most cases, our approach performs superior in terms of average squared error between the estimated and true time-varying spectral density. Finally, we demonstrate our methodology through applications to quarterly Gross Domestic Product (GDP) data and Northern California wind data.

stat.ME