SearcharxivSearch

arXiv subjects

Sho Kawano

Publications and source records attributed to Sho Kawano.

3 recordsLinked to original sources

Nonprobability Samples for Small Area Estimation: A Review and Comparative Simulation Study

Nonprobability samples (NPS) are attractive because they are less costly to collect, can provide substantially larger sample sizes, and may reach populations that traditional probability surveys do not. As response rates for traditional surveys fall, interest in NPS has grown rapidly within the field of survey statistics. These methods are especially relevant for small area estimation (SAE), where there is ever-present demand for estimates at fine geographic scales and detailed demographic domains. Despite rapid methodological development, there remains limited understanding of which approaches perform best under different conditions. In this paper, we review recent developments in NPS methodology, including the concept of data defect correlation (DDC) as a measure of data quality and as a tool for categorizing the various NPS methods. We then present a comprehensive simulation study that evaluates a range of NPS approaches under varying levels of DDC and extend several existing methods to the SAE setting.

stat.ME

On Data Thinning for Model Validation in Small Area Estimation

Small area estimation produces estimates of population parameters for geographic and demographic subgroups with limited sample sizes. Such estimates are critical for policy decisions, yet principled validation of these models remains a challenge. Unlike conventional predictive settings, validation data are rarely available. Data thinning splits a single observation into independent training and test components. It enables out-of-sample validation using only the area-level summary statistics routinely available, requiring only their Gaussianity and known sampling variances. However, the properties of thinning-based model comparison have not been formally studied. In this paper, we develop these properties. We construct an unbiased estimator of thinned-data mean squared error and show that it differs systematically from its full-data counterpart; for the standard Fay-Herriot model, the gap admits a closed-form expression that depends on the candidate model's shrinkage behavior. We further show that the estimator variance increases sharply as the training fraction approaches one, producing a bias-variance tradeoff with no universally optimal thinning parameter. Practical recommendations balancing these forces are informed by theory and verified empirically. Design-based simulations using American Community Survey microdata show that the recommended data thinning approach is competitive with information-criterion and simulation-based methods, and substantially more stable across heterogeneous sampling designs.

stat.ME

Spatially Selected and Dependent Random Effects for Small Area Estimation with Application to Rent Burden

Area-level models for small area estimation typically rely on areal random effects to shrink design-based direct estimates towards a model-based predictor. Incorporating the spatial dependence of the random effects into these models can further improve the estimates when there are not enough covariates to fully account for spatial dependence of the areal means. A number of recent works have investigated models that include random effects for only a subset of areas, in order to improve the precision of estimates. However, such models do not readily handle spatial dependence. In this paper, we introduce a model that accounts for spatial dependence in both the random effects as well as the latent process that selects the effects. We show how this model can significantly improve predictive accuracy via an empirical simulation study based on data from the American Community Survey, and illustrate its properties via an application to estimate county-level median rent burden.

stat.ME