SearcharxivSearch

arXiv subjects

Daniel Vedensky

Publications and source records attributed to Daniel Vedensky.

4 recordsLinked to original sources

Nonprobability Samples for Small Area Estimation: A Review and Comparative Simulation Study

Nonprobability samples (NPS) are attractive because they are less costly to collect, can provide substantially larger sample sizes, and may reach populations that traditional probability surveys do not. As response rates for traditional surveys fall, interest in NPS has grown rapidly within the field of survey statistics. These methods are especially relevant for small area estimation (SAE), where there is ever-present demand for estimates at fine geographic scales and detailed demographic domains. Despite rapid methodological development, there remains limited understanding of which approaches perform best under different conditions. In this paper, we review recent developments in NPS methodology, including the concept of data defect correlation (DDC) as a measure of data quality and as a tool for categorizing the various NPS methods. We then present a comprehensive simulation study that evaluates a range of NPS approaches under varying levels of DDC and extend several existing methods to the SAE setting.

stat.ME

Bayesian Unit-level Modeling of Categorical Survey Data with a Longitudinal Design

Categorical response data are ubiquitous in complex survey applications, yet few methods model the dependence across different outcome categories when the response is ordinal. Likewise, few methods exist for the common combination of a longitudinal design and categorical data. By modeling individual survey responses at the unit-level, it is possible to capture both ordering information in ordinal responses and any longitudinal correlation. However, accounting for a complex survey design becomes more challenging in the unit-level setting. We propose a Bayesian hierarchical, unit-level, model-based approach for categorical data that is able to capture ordering among response categories, can incorporate longitudinal dependence, and accounts for the survey design. To handle computational scalability, we develop efficient Gibbs samplers with appropriate data augmentation as well as variational Bayes algorithms. Using public-use microdata from the Household Pulse Survey, we provide an analysis of an ordinal response that asks about the frequency of anxiety symptoms at the beginning of the COVID-19 pandemic. We compare both design-based and model-based estimators and demonstrate superior performance for the proposed approaches.

stat.ME

Bayesian Unit-level Models for Longitudinal Survey Data under Informative Sampling: An Analysis of Expected Job Loss Using the Household Pulse Survey

The Household Pulse Survey (HPS), recently released by the U.S. Census Bureau, gathers timely information about the societal and economic impacts of coronavirus. The first phase of the survey was quickly launched one month after the beginning of the coronavirus pandemic and ran for 12 weeks. To track the immediate impact of the pandemic, individual respondents during this phase were re-sampled for up to three consecutive weeks. Motivated by expected job loss during the pandemic, using public-use microdata, this work proposes unit-level, model-based estimators that incorporate longitudinal dependence at both the response and domain level. In particular, using a pseudo-likelihood, we consider a Bayesian hierarchical unit-level, model-based approach for both Gaussian and binary response data under informative sampling. To facilitate construction of these model-based estimates, we develop an efficient Gibbs sampler. An empirical simulation study is conducted to compare the proposed approach to models that do not account for unit-level longitudinal correlation. Finally, using public-use HPS micro-data, we provide an analysis of "expected job loss" that compares both design-based and model-based estimators and demonstrates superior performance for the proposed model-based approaches.

stat.ME

A Look into the Problem of Preferential Sampling from the Lens of Survey Statistics

An evolving problem in the field of spatial and ecological statistics is that of preferential sampling, where biases may be present due to a relationship between sample data locations and a response of interest. This field of research bears a striking resemblance to the longstanding problem of informative sampling within survey methodology, although with some important distinctions. With the goal of promoting collaborative effort within and between these two problem domains, we make comparisons and contrasts between the two problem statements. Specifically, we review many of the solutions available to address each of these problems, noting the important differences in modeling techniques. Additionally, we construct a series of simulation studies to examine some of the methods available for preferential sampling, as well as a comparison analyzing heavy metal biomonitoring data.

stat.ME