SearcharxivSearch

arXiv subjects

Grayson W. White

Publications and source records attributed to Grayson W. White.

5 recordsLinked to original sources

Differentially Private Average Treatment Effect Estimation by Propensity Score Blocking

Average treatment effect (ATE) estimation in observational studies is a fundamental statistical tool used frequently in social science, medicine, and other fields. These fields often work with sensitive data where privacy protections are important, so a differentially private mechanism for ATE estimation is highly desirable. Here we present two propensity score-based algorithms for ATE estimation on observational data, one improving the inverse probability weighting (IPW) method used in prior work, and the other using blocking on the propensity score (BPS). Both show lower error and less bias than prior work, with the BPS-based algorithm frequently reducing error by 75% or more compared to prior work.

stat.ME

A penalized logistic generalized regression estimator

Under a model-assisted framework, a penalized logistic generalized regression estimator is developed to estimate a finite population proportion from complex survey data and auxiliary data. The proposed estimator controls the impact of unnecessary auxiliary variables through a lasso or ridge penalty. A central limit theorem is derived for the penalized regression coefficients and for the penalized logistic generalized regression estimator. Through simulations, it is shown that including the penalty increases the efficiency of the estimator when the true model is sparse. An example using United States Forest Service data demonstrates the applicability of the estimator for surveys with many available auxiliary variables.

stat.ME

Hierarchical models for small area estimation using zero-inflated forest inventory variables: comparison and implementation

National Forest Inventory (NFI) data are typically limited to sparse networks of sample locations due to cost constraints. While design-based estimators provide reliable forest parameter estimates for large areas, there is increasing interest in model-based small area estimation (SAE) methods to improve precision for smaller spatial, temporal, or biophysical domains. SAE methods can be broadly categorized into area- and unit-level models, with unit-level models offering greater flexibility, making them the focus of this study. Ensuring valid inference requires satisfying model distributional assumptions, which is particularly challenging for NFI variables that exhibit positive support and zero-inflation, such as forest biomass, carbon, and volume. Here, we evaluate nine candidate estimators, including two-stage unit-level hierarchical Bayesian models, single-stage Bayesian models, and two-stage frequentist models, for estimating forest biomass at the county level in Nevada and Washington, United States. Estimator performance is assessed using repeated sampling from simulated populations and unit-level cross-validation with FIA data. Results show that small area estimators incorporating a two-stage approach to account for zero-inflation, county-specific random intercepts and residual variances, and spatial random effects yield the most accurate and well-calibrated county-level estimates, with spatial effects providing the greatest benefits when spatial autocorrelation is present in the underlying population.

stat.AP

A method for empirically assessing small area estimators via bootstrap-weighted k-Nearest-Neighbor artificial populations, with applications to forest inventory

National Forest Inventories (NFIs) monitor forest attributes across a variety of spatial and temporal scales in a given country. Increased interest in reporting and management at smaller scales has driven NFIs to investigate and adopt small area estimation (SAE) due to the promise of increased precision at these scales. However, comparing and evaluating SAE models for a given application is inherently difficult. Typically, many areas lack enough data to check unit-level modeling assumptions or to assess unit-level predictions empirically; and no ground truth is available for checking area-level estimates. Design-based simulation from artificial populations can help with each of these issues, but only if the artificial populations realistically represent the application at hand and are not built using assumptions that inherently favor one SAE model over another. In this paper, we borrow ideas from random hot deck, approximate Bayesian bootstrap (ABB), and k Nearest Neighbor (kNN) imputation methods to propose a kNN-based approximation to ABB (KBAABB), for generating an artificial population when rich unit-level auxiliary data is available. We introduce diagnostic checks on the process of building the artificial population, and we demonstrate how to use such an artificial population for design-based simulation studies to compare and evaluate SAE models, using real data from the United States Department of Agriculture, Forest Service, Forest Inventory and Analysis Program, the NFI of the United States.

stat.ME

Small area estimation of forest biomass via a two-stage model for continuous zero-inflated data

The United States (US) Forest Inventory & Analysis Program (FIA) collects data on and monitors the trends of forests in the US. FIA is increasingly interested in monitoring forest attributes such as biomass at fine geographic and temporal scales, resulting in a need for assessment and development of small area estimation techniques in forest inventory. We implement a small area estimator and parametric bootstrap estimator that account for zero-inflation in biomass data via a two-stage model-based approach and compare its performance to a post-stratified estimator and to the unit- and area-level empirical best linear unbiased prediction (EBLUP) estimators. For estimator comparison, we conduct a simulation study with counties in the US state Nevada as domains based on sampled plot data and remote sensing data products. Results show the zero-inflated estimator has the lowest relative bias and the smallest empirical root mean square error. Moreover, the 95% confidence interval coverages of the zero-inflated estimator and the unit-level EBLUP are more accurate than the other two estimators. To further illustrate the practical utility, we employ a data application across the 2019 measurement year in Nevada. We introduce the R package, saeczi, which efficiently implements the zero-inflated estimator and its mean squared error estimator.

stat.AP