SearcharxivSearch

arXiv subjects

Swen Kuh

Publications and source records attributed to Swen Kuh.

3 recordsLinked to original sources

When Bayes goes bad: Weakly-regularized covariate adjustment leads to a biased estimate of prevalence

When estimating population prevalence from a non-random sample, it is important to adjust for differences between sample and population. However, adjustment for multiple factors requires analysis that can be difficult to understand and validate. In this manuscript, we explore an unexpected downward trend of estimates when covariates are added sequentially to a Bayesian hierarchical model for the estimation of the prevalence of SARS-CoV-2 specific antibodies in an Australian city in late 2020. We compare our data analysis to results from a simulation study to understand four potential contributors to this effect: (i) correction for differences between sample and population, (ii) rare-events bias in logistic regression, (iii) inclusion of the uncertainty of test sensitivity and specificity in a multilevel model, and (iv) increasing model dimensionality. We find that weak prior distributions on the logistic regression coefficients lead to a systematic increase in the amount of partial pooling across adjustment cells-the prior becomes stronger as model dimensionality increases-which in turn feeds through to the estimated assay specificity, which then feeds back to the model and results in lowering the estimated prevalence. Our paper contributes three elements: (i) immediate and longer-term recommendations for using these types of models, (ii) simulation studies to explore the impact of the contributors to this effect, and (iii) a worked example of investigation of unexpected results in a model with multiple adjustment factors.

stat.ME

Latent Causal Socioeconomic Health Index

This research develops a model-based LAtent Causal Socioeconomic Health (LACSH) index at the national level. Motivated by the need for a holistic national well-being index, we build upon the latent health factor index (LHFI) approach that has been used to assess the unobservable ecological/ecosystem health. LHFI integratively models the relationship between metrics, latent health, and covariates that drive the notion of health. In this paper, the LHFI structure is integrated with spatial modeling and statistical causal modeling. Our efforts are focused on developing the integrated framework to facilitate the understanding of how an observational continuous variable might have causally affected a latent trait that exhibits spatial correlation. A novel visualization technique to evaluate covariate balance is also introduced for the case of a continuous policy (treatment) variable. Our resulting LACSH framework and visualization tool are illustrated through two global case studies on national socioeconomic health (latent trait), each with various metrics and covariates pertaining to different aspects of societal health, and the treatment variable being mandatory maternity leave days and government expenditure on healthcare, respectively. We validate our model by two simulation studies. All approaches are structured in a Bayesian hierarchical framework and results are obtained by Markov chain Monte Carlo techniques.

stat.ME

Using leave-one-out cross-validation (LOO) in a multilevel regression and poststratification (MRP) workflow: A cautionary tale

In recent decades, multilevel regression and poststratification (MRP) has surged in popularity for population inference. However, the validity of the estimates can depend on details of the model, and there is currently little research on validation. We explore how leave-one-out cross-validation (LOO) can be used to compare Bayesian models for MRP. We investigate two approximate calculations of LOO, the Pareto smoothed importance sampling (PSIS-LOO) and a survey-weighted alternative (WTD-PSIS-LOO). Using two simulation designs, we examine how accurately these two criteria recover the correct ordering of model goodness at predicting population and small area level estimands. Focusing first on variable selection, we find that neither PSIS-LOO nor WTD-PSIS-LOO correctly recovers the models' order for an MRP population estimand (although both criteria correctly identify the best and worst model). When considering small-area estimation, the best model differs for different small areas, highlighting the complexity of MRP validation. When considering different priors, the models' order seems slightly better at smaller area levels. These findings suggest that while not terrible, PSIS-LOO-based ranking techniques may not be suitable to evaluate MRP as a method. We suggest this is due to the aggregation stage of MRP, where individual-level prediction errors average out. These results show that in practice, PSIS-LOO-based model validation tools need to be used with caution and might not convey the full story when validating MRP as a method.

stat.ME