SearcharxivSearch

arXiv subjects

Jon Wakefield

Publications and source records attributed to Jon Wakefield.

At least 19 recordsLinked to original sources

Sampling distributions for complex design variance estimators in a Fay-Herriot model

Fay-Herriot (FH) models with variance smoothing typically use chi-squared sampling distributions for the design variance estimators. This choice is only valid under strong assumptions on the population and the sampling design, and the choice of sampling distribution is understudied for complex survey designs such as the stratified two-stage clustering design used by the Demographic and Health Surveys (DHS). DHS conducts surveys in low- and middle-income countries and result in low sample sizes for unplanned domains of interest. Thus, accounting for the uncertainty in the estimated design variances is important. We derive two sampling distributions under the DHS design, a simple and a more complex, while clearly specifying and discussing the required superpopulation and design assumptions. In a simulation study, we compare the two sampling distributions to the empirical sampling distributions, and the resulting FH models with variance smoothing to the standard FH model. We find that the standard model exhibits undercoverage, while the variance smoothing models produce better credible intervals according to proper scoring rules. Interestingly, the simple sampling distribution, which is easiest to implement, performs equally as well as the more complex sampling distribution. We illustrate the proposed models by estimating height-for-age z-scores using the 2022 Kenya DHS.

stat.ME

Automatic Variance Adjustment for Small Area Estimation

Small area estimation (SAE) is a common endeavor and is used in a variety of disciplines. In low- and middle-income countries (LMICs), in which household surveys provide the most reliable and timely source of data, SAE is vital for highlighting disparities in health and demographic indicators. Weighted estimators are ideal for inference, but for fine geographical partitions in which there are insufficient data, SAE models are required. The most common approach is Fay-Herriot area-level modeling in which the data requirements are a weighted estimate and an associated variance estimate. The latter can be undefined or unstable when data are sparse and so we propose a principled modification which is based on augmenting the available data with a prior sample from a hypothetical survey. This adjustment is generally available, respects the design and is simple to implement. We examine the empirical properties of the adjustment through simulation and illustrate its use with wasting data from a 2018 Zambian Demographic and Health Survey. The modification is implemented as an automatic remedy in the R package surveyPrev, which provides a comprehensive suite of tools for conducing SAE in LMICs.

stat.ME

A Survival Framework for Estimating Child Mortality Rates using Multiple Data Types

Child mortality is an important population health indicator. However, many countries lack high-quality vital registration to measure child mortality rates precisely and reliably over time. Research endeavors such as those by the United Nations Inter-agency Group for Child Mortality Estimation (UN IGME) and the Global Burden of Disease (GBD) study leverage statistical models and available data to estimate child survival summaries including neonatal, infant, and under-five mortality rates. UN IGME fits separate models for each age group and the GBD uses a multi-step modeling process. We propose a Bayesian survival framework to estimate temporal trends in the probability of survival as a function of age, up to the fifth birthday, with a single model. Our framework integrates all data types that are used by UN IGME: household surveys, vital registration, and other pre-processed mortality rates. We demonstrate that our framework is applicable to any country using log-logistic and piecewise-exponential survival functions, and discuss findings for four example countries with diverse data profiles: Kenya, Brazil, Estonia, and Syrian Arab Republic. Our model produces estimates of the three survival summaries that are in broad agreement with both the data and the UN IGME estimates, but in addition gives the complete survival curve.

stat.AP

Small Area Estimation Methods for Multivariate Health and Demographic Outcomes using Complex Survey Data

Improving health in the most disadvantaged populations requires reliable estimates of health and demographic indicators to inform policy and interventions. Low- and middle-income countries with the largest burden of disease and disability tend to have the least comprehensive data, relying primarily on household surveys. Subnational estimates are increasingly used to inform targeted interventions and health policies. Producing reliable estimates from these data at fine geographical scales requires statistical modeling, and small area estimation models are commonly used in this context. Although most current methods model univariate outcomes, improved estimates may be attained by borrowing strength across related outcomes via multivariate modeling. In this paper, we develop classes of area- and unit-level multivariate shared component models using complex survey data. This framework jointly models multiple outcomes to improve accuracy of estimates compared to separately fitting univariate models. We conduct simulation studies to validate the methodology and use the proposed approach on survey data from Kenya in 2014; first, to jointly model height-for-age and weight-for-age in children, and second, to model three categories of contraceptive use in women. These models produce improved estimates compared to univariate and naive multivariate modeling approaches.

stat.ME

Small Area Estimation of Fertility in Low- and Middle-Income Countries

Accurate fertility estimates at fine spatial resolution are essential for localized public health planning, particularly in low- and middle-income countries (LMICs). While national-level indicators such as age-specific fertility rates (ASFR) and total fertility rate (TFR) are often reported through official statistics, they lack the spatial granularity needed to guide targeted interventions. To address this, we develop a framework for subnational fertility estimation using small-area estimation (SAE) techniques applied to birth history data from household surveys, in particular Demographic and Health Surveys (DHS). Disaggregation by geographic area, time period, and maternal age group leads to significant data sparsity, limiting the reliability of direct estimates at fine scales. To overcome this, we propose a suite of methods, including direct estimators, area-level and unit-level Bayesian hierarchical models, to produce accurate estimates across varying spatial resolutions. The model-based approaches incorporate spatiotemporal smoothing and integrate covariates such as maternal education, contraceptive use and urbanicity. Using data from the 2021 Madagascar DHS, we generate district-level ASFR and TFR estimates and evaluate model performance through cross-validation.

stat.ME

sae4health: An R Shiny Application for Small Area Estimation in Low- and Middle-Income Countries

Accurate subnational estimation of health indicators is critical for public health planning, particularly in low- and middle-income countries (LMICs), where data and analytic tools are often limited. sae4health is an open-access Shiny application (https://rsc.stat.washington.edu/sae4health/) that generates small area estimates for more than 150 demographic and health indicators, based on over 150 Demographic and Health Surveys (DHS) from 60 countries. The platform offers both area- and unit-level models with spatial random effects, implemented through fast Bayesian inference using Integrated Nested Laplace Approximation (INLA). The app is fully browser-based and requires no data input, programming skills, or statistical modeling expertise, making advanced methods accessible to a wide range of users. Estimates are processed in real time and presented as interactive maps, tables, and downloadable reports. A companion website (https://sae4health.stat.uw.edu) provides documentation and methodological background to support the app. Together, these resources enhance access to subnational health data and facilitate the use of DHS surveys for evidence-based decision making.

stat.CO

Toward a Principled Workflow for Prevalence Mapping Using Household Survey Data

Understanding the prevalence of key demographic and health indicators in small geographic areas and domains is of global interest, especially in low- and middle-income countries (LMICs), where vital registration data is sparse and household surveys are the primary source of information. Recent advances in computation and the increasing availability of spatially detailed datasets have led to much progress in sophisticated statistical modeling of prevalence. As a result, high-resolution prevalence maps for many indicators are routinely produced in the literature. However, statistical and practical guidance for producing prevalence maps in LMICs has been largely lacking. In particular, advice in choosing and evaluating models and interpreting results is needed, especially when data is limited. Software and analysis tools are also usually inaccessible to researchers in low-resource settings to conduct their own analysis or reproduce findings in the literature. In this paper, we propose a general workflow for prevalence mapping using household survey data. We consider all stages of the analysis pipeline, with particular emphasis on model choice and interpretation. We illustrate the proposed workflow using a case study mapping the proportion of pregnant women who had at least four antenatal care visits in Kenya. Reproducible code is provided in the Supplementary Materials and can be readily extended to a broad collection of indicators.

stat.AP

Evaluating Multilevel Regression and Poststratification with Spatial Priors with a Big Data Behavioural Survey

Multilevel regression and poststratification (MRP) is a computationally efficient indirect estimation method that can quickly produce improved population-adjusted estimates with limited data. Recent computational advancements allow efficient, relatively simple, and quick approximate Bayesian estimation for MRP. As population health outcomes of interest including vaccination uptake are known to have spatial structure, precision may be gained by including space in the model. We test a recently proposed spatial MRP method that includes a BYM2 spatial term that smooths across demographics and geographic areas using a large, unrepresentative survey. We produce California county-level estimates of first-dose COVID-19 vaccination up to June 2021 using classic and spatial MRP models, and poststratify using data from the American Community Survey (US Census Bureau). We assess validity using reported first-dose vaccination counts from the Centers for Disease Control (CDC). Neither classic nor spatial MRP models performed well, highlighting: 1. spatial MRP may be most appropriate for richer data contexts, 2. some demographics in the survey data are over-sampled and -aggregated, producing model over-smoothing, and 3. a need for survey producers to share user-representative metrics to better benchmark estimates.

stat.AP

Exact variance estimation for model-assisted survey estimators using U- and V-statistics

Model-assisted estimation combines sample survey data with auxiliary information to increase precision when estimating finite population quantities. Accurately estimating the variance of model-assisted estimators is challenging: the classical approach ignores uncertainty from estimating the working model for the functional relationship between survey and auxiliary variables. This approach may be asymptotically valid, but can underestimate variance in practical settings with limited sample sizes. In this work, we develop a connection between model-assisted estimation and the theory of U- and V-statistics. We demonstrate that when predictions from the working model for the variable of interest can be represented as a U- or V-statistic, the resulting model-assisted estimator also admits a U- or V-statistic representation. We exploit this connection to derive an improved estimator of the exact variance of such model-assisted estimators. The class of working models for which this strategy can be used is broad, ranging from linear models to modern ensemble methods. We apply our approach to the model-assisted estimator constructed with a linear regression working model, commonly referred to as the generalized regression estimator, show that it can be re-written as a U-statistic, and propose an estimator of its exact variance. We illustrate our proposal and compare it against the classical asymptotic variance estimator using household survey data from the American Community Survey.

stat.ME

Small Area Estimation of Education Levels in Low- and Middle-Income Countries

Education is a key driver of social and economic mobility, yet disparities in attainment persist, particularly in low- and middle-income countries (LMICs). Existing indicators, such as mean years of schooling for adults aged 25 and older (MYS25) and expected years of schooling (EYS), offer a snapshot of an educational system, but lack either cohort-specific or temporal granularity. To address these limitations, we introduce the ultimate years of schooling (UYS)-a birth cohort-based metric targeting the final educational attainment of any individual cohort, including those with ongoing schooling trajectories. As with many attainment indicators, we propose to estimate UYS with cross-sectional household surveys. However, for younger cohorts, estimation fails, because these individuals are right-censored leading to severe downwards bias. To correct for this, we propose to re-frame educational attainment as a time-to-event process and deploy discrete-time survival models that explicitly account for censoring in the observations. At the national level, we estimate the parameters of the model using survey-weighted logistic regression, while for finer spatial resolutions, where sample sizes are smaller, we embed the discrete-time survival model within a Bayesian spatiotemporal framework to improve stability and precision. Applying our proposed methods to data from the 2022 Tanzania Demographic and Health Surveys, we estimate female educational trajectories corrected for censoring biases, and reveal substantial subnational disparities. By providing a dynamic, bias-corrected, and spatially disaggregated measure, our approach enhances education monitoring; it equips policymakers and researchers with a more precise tool for monitoring current progress towards education goals, and for designing future targeted policy interventions in LMICs.

stat.ME

Direct-Assisted Bayesian Unit-level Modeling for Small Area Estimation of Rare Event Prevalence

Small area estimation using survey data can be achieved by using either a design-based or a model-based inferential approach. Design-based direct estimators are generally preferable because of their consistency, asymptotic normality, and reliance on fewer assumptions. However, when data are sparse at the desired area level, as is often the case when measuring rare events, these direct estimators can have extremely large uncertainty, making a model-based approach preferable. A model-based approach with a random spatial effect borrows information from surrounding areas at the cost of inducing shrinkage. As a result, estimates may be over-smoothed and inconsistent with design-based estimates at higher area levels when aggregated. We propose two unit-level Bayesian models for small area estimation of rare event prevalence which use design-based direct estimates at a higher area level to increase consistency in aggregation. This model framework is designed to accommodate sparse data obtained from two-stage stratified cluster sampling, which is particularly relevant to applications in low- and middle-income countries. After introducing the model framework and its implementation, we conduct a simulation study to evaluate its properties and apply it to the estimation of the neonatal mortality rate in Zambia, using 2014 Demographic Health Surveys data.

stat.ME

A Pseudo-likelihood Approach to Under-5 Mortality Estimation

Accurate and precise estimates of the under-5 mortality rate (U5MR) are an important health summary for countries. However, full survival curves allow us to better understand the pattern of mortality in children under five. Modern demographic methods for estimating a full mortality schedule for children have been developed for countries with good vital registration and reliable census data, but perform poorly in many low- and middle-income countries (LMICs). In these countries, the need to utilize nationally representative surveys to estimate the U5MR requires additional care to mitigate potential biases in survey data, acknowledge the survey design, and handle the usual characteristics of survival data, for example, censoring and truncation. In this paper, we develop parametric and non-parametric pseudo-likelihood approaches to estimating child mortality across calendar time from complex survey data. We show that the parametric approach is particularly useful in scenarios where data are sparse and parsimonious models allow efficient estimation. We compare a variety of parametric models to two existing methods for obtaining a full survival curve for children under the age of 5, and argue that a parametric pseudo-likelihood approach is advantageous in LMICs. We apply our proposed approaches to survey data from four LMICs.

stat.ME

Incorporating testing volume into estimation of effective reproduction number dynamics

Branching process inspired models are widely used to estimate the effective reproduction number -- a useful summary statistic describing an infectious disease outbreak -- using counts of new cases. Case data is a real-time indicator of changes in the reproduction number, but is challenging to work with because cases fluctuate due to factors unrelated to the number of new infections. We develop a new model that incorporates the number of diagnostic tests as a surveillance model covariate. Using simulated data and data from the SARS-CoV-2 pandemic in California, we demonstrate that incorporating tests leads to improved performance over the state-of-the-art.

stat.ME

BARTSIMP: flexible spatial covariate modeling and prediction using Bayesian additive regression trees

Prediction is a classic challenge in spatial statistics and the inclusion of spatial covariates can greatly improve predictive performance when incorporated into a model with latent spatial effects. It is desirable to develop flexible regression models that allow for nonlinearities and interactions in the covariate specification. Existing machine learning approaches that allow for spatial dependence in the residuals fail to provide reliable uncertainty estimates. In this paper, we investigate the combination of a Gaussian process spatial model with a Bayesian Additive Regression Tree (BART) model. The computational burden of the approach is reduced by combining Markov chain Monte Carlo (MCMC) with the Integrated Nested Laplace Approximation (INLA) technique. We study the performance of the method first via simulation. We then use the model to predict anthropometric responses in Kenya, with the data collected via a complex sampling design. In particular, household survey data are collected via stratified two-stage unequal probability cluster sampling, which requires special care when modeled.

stat.ME

Pseudo-Bayesian unit level modeling for small area estimation under informative sampling

When mapping subnational health and demographic indicators, direct weighted estimators of small area means based on household survey data can be unreliable when data are limited. If survey microdata are available, unit level models can relate individual survey responses to unit level auxiliary covariates and explicitly account for spatial dependence and between area variation using random effects. These models can produce estimators with improved precision, but often neglect to account for the design of the surveys used to collect data. Pseudo-Bayesian approaches incorporate sampling weights to address informative sampling when using such models to conduct population inference but credible sets based on the resulting pseudo-posterior distributions can be poorly calibrated without adjustment. We outline a pseudo-Bayesian strategy for small area estimation that addresses informative sampling and incorporates a post-processing rescaling step that produces credible sets with close to nominal empirical frequentist coverage rates. We compare our approach with existing design-based and model-based estimators using real and simulated data.

stat.ME

Estimating Subnational Under-Five Mortality Rates Using a Spatio-Temporal Age-Period-Cohort Model

Producing subnational estimates of the under-five mortality rate (U5MR) is a vital goal for the United Nations to reduce inequalities in mortality and well-being across the globe. There is a great disparity in U5MR between high-income and Low-and-Middle Income Countries (LMICs). Current methods for modelling U5MR in LMICs use smoothing methods to reduce uncertainty in estimates caused by data sparsity. This paper includes cohort alongside age and period in a novel application of an Age-Period-Cohort model for U5MR. In this context, current methods only use age and period (and not cohort) for smoothing. With data from the Kenyan Demographic and Health Surveys (DHS) we use a Bayesian hierarchical model with terms to smooth over temporal and spatial components whilst fully accounting for the complex stratified, multi-staged cluster design of the DHS. Our results show that the use of cohort may be useful in the context of subnational estimates of U5MR. We validate our results at the subnational level by comparing our results against direct estimates.

stat.AP

Adaptive Gaussian Markov Random Fields for Child Mortality Estimation

The under-5 mortality rate (U5MR), a critical health indicator, is typically estimated from household surveys in lower and middle income countries. Spatio-temporal disaggregation of household survey data can lead to highly variable estimates of U5MR, necessitating the usage of smoothing models which borrow information across space and time. The assumptions of common smoothing models may be unrealistic when certain time periods or regions are expected to have shocks in mortality relative to their neighbors, which can lead to oversmoothing of U5MR estimates. In this paper, we develop a spatial and temporal smoothing approach based on Gaussian Markov random field models which incorporate knowledge of these expected shocks in mortality. We demonstrate the potential for these models to improve upon alternatives not incorporating knowledge of expected shocks in a simulation study. We apply these models to estimate U5MR in Rwanda at the national level from 1985-2019, a time period which includes the Rwandan civil war and genocide.

stat.ME

Small Area Estimation with Random Forests and the LASSO

We consider random forests and LASSO methods for model-based small area estimation when the number of areas with sampled data is a small fraction of the total areas for which estimates are required. Abundant auxiliary information is available for the sampled areas, from the survey, and for all areas, from an exterior source, and the goal is to use auxiliary variables to predict the outcome of interest. We compare areal-level random forests and LASSO approaches to a frequentist forward variable selection approach and a Bayesian shrinkage method. Further, to measure the uncertainty of estimates obtained from random forests and the LASSO, we propose a modification of the split conformal procedure that relaxes the assumption of identically distributed data. This work is motivated by Ghanaian data available from the sixth Living Standard Survey (GLSS) and the 2010 Population and Housing Census. We estimate the areal mean household log consumption using both datasets. The outcome variable is measured only in the GLSS for 3\% of all the areas (136 out of 5019) and more than 170 potential covariates are available from both datasets. Among the four modelling methods considered, the Bayesian shrinkage performed the best in terms of bias, MSE and prediction interval coverages and scores, as assessed through a cross-validation study. We find substantial between-area variation, the log consumption areal point estimates showing a 1.3-fold variation across the GAMA region. The western areas are the poorest while the Accra Metropolitan Area district gathers the richest areas.

stat.ME