SearcharxivSearch

arXiv subjects

Kevin P. Josey

Publications and source records attributed to Kevin P. Josey.

10 recordsLinked to original sources

Causal Inference with Multiple Misclassified Exposures: A Control Variate-Adjusted Calibration Weighting Approach

Exposure misclassification is a common issue prevalent in studies of respiratory infections in cystic fibrosis. Throat swabs are frequently used in place of expectorated or induced sputum cultures, although they have imperfect sensitivity and specificity to detect Pseudomonas aeruginosa and Staphylococcus aureus. We develop calibration weighting and control variate estimators for causal inference with multiple misclassified binary exposures and clustered observations. The calibration approach treats misclassification as a missing data problem, achieving consistency without modeling the misclassification mechanism. The control variate adjustment integrates information from error-prone observations to reduce variance while preserving the consistency of the gold-standard estimator. We show that the resulting estimator inherits double robustness from its component estimators. We also characterize a structural ceiling on efficiency gains in the bivariate setting, where joint correct classification of both exposures limits the variance reduction achievable relative to univariate applications. Simulation studies confirm the consistency and double robustness of the proposed estimators under model misspecification. We then apply these methods to a cohort of $651$ cystic fibrosis patients ages $6$--$21$. Swab-based estimates attenuate the effect of P. aeruginosa on percent predicted FEV$_1$ by approximately $79\%$ relative to sputum-based estimates ($-1.68$ vs. $-7.88$ percentage points; $95\%$ confidence interval for sputum: $-10.08$, $-5.67$). These findings suggest that relying on throat swabs may lead to under-treatment of P. aeruginosa infections. More broadly, the methods provide a framework for causal inference with multiple misclassified exposures.

stat.ME

Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding

Data-driven decision making frequently relies on predicting counterfactual outcomes. In practice, researchers commonly train counterfactual prediction models on a source dataset to inform decisions on a possibly separate target population. Conformal prediction has arisen as a popular method for producing assumption-lean prediction intervals for counterfactual outcomes that would arise under different treatment decisions in the target population of interest. However, existing methods require that every confounding factor of the treatment-outcome relationship used for training on the source data is additionally measured in the target population, risking miscoverage if important confounders are unmeasured in the target population. In this paper, we introduce a computationally efficient debiased machine learning framework that allows for valid prediction intervals when only a subset of confounders is measured in the target population, a common challenge referred to as runtime confounding. Grounded in semiparametric efficiency theory, we show the resulting prediction intervals achieve desired coverage rates with faster convergence compared to standard methods. Through numerous synthetic and semi-synthetic experiments, we demonstrate the utility of our proposed method.

stat.ML

Transporting Predictions via Double Machine Learning: Predicting Partially Unobserved Students' Outcomes

Educational policymakers often lack data on student outcomes where standardized tests were not administered. Machine learning can predict unobserved outcomes in target populations using source population data. However, covariate distribution differences between populations reduce model transportability, potentially decreasing predictive accuracy and introducing bias. We propose using double machine learning for covariate-shift weighted models. First, we estimate overlap scores -- the probability an observation belongs to the source dataset given covariates. Second, balancing weights, defined as density ratios of target-to-source membership probabilities, reweight individual observations' contributions to the loss function in target outcome prediction models. This downweights source observations less similar to the target population, allowing predictions to rely more on observations with greater overlap. Consequently, predictions become more transportable under covariate shift. We illustrate this framework using student standardized financial literacy scores (FLS) data. Using Bayesian Additive Regression Trees (BART), we predict missing FLS. We find minimal predictive performance differences between weighted and unweighted models, suggesting limited covariate shift in our setting. Nonetheless, our approach provides a principled framework for addressing covariate shift and is broadly applicable to predictive modeling in social and health sciences, where source-target population differences are common.

stat.AP

Flexible and Efficient Estimation of Causal Effects with Error-Prone Exposures: A Control Variates Approach for Measurement Error

Exposure measurement error is a ubiquitous but often overlooked challenge in causal inference with observational data. Existing methods accounting for exposure measurement error largely rely on restrictive parametric assumptions, while emerging data-adaptive estimation approaches allow for less restrictive assumptions but at the cost of flexibility, as they are typically tailored towards rigidly-defined statistical quantities. There remains a critical need for assumption-lean estimation methods that are both flexible and possess desirable theoretical properties across a variety of study designs. In this paper, we introduce a general framework for estimation of causal quantities in the presence of exposure measurement error, adapted from the control variates approach of Yang and Ding (2019). Our method can be implemented in various two-phase sampling study designs, where one obtains gold-standard exposure measurements for a small subset of the full study sample, called the validation data. The control variates framework leverages both the error-prone and error-free exposure measurements by augmenting an initial consistent estimator from the validation data with a variance reduction term formed from the full data. We show that our method inherits double-robustness properties under standard causal assumptions. Simulation studies show that our approach performs favorably compared to leading methods under various two-phase sampling schemes. We illustrate our method with observational electronic health record data on HIV outcomes from the Vanderbilt Comprehensive Care Clinic.

stat.ME

Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements

Measurement error is a common challenge for causal inference studies using electronic health record (EHR) data, where clinical outcomes and treatments are frequently mismeasured. Researchers often address measurement error by conducting manual chart reviews to validate measurements in a subset of the full EHR data -- a form of two-phase sampling. To improve efficiency, phase-two samples are often collected in a biased manner dependent on the patients' initial, error-prone measurements. In this work, motivated by our aim of performing causal inference with error-prone outcome and treatment measurements under two-phase sampling, we develop solutions applicable to both this specific problem and the broader problem of causal inference with two-phase samples. For our specific measurement error problem, we construct two asymptotically equivalent doubly-robust estimators of the average treatment effect and demonstrate how these estimators arise from two previously disconnected approaches to constructing efficient estimators in general two-phase sampling settings. We document various sources of instability affecting estimators from each approach and propose modifications that can considerably improve finite sample performance in any two-phase sampling context. We demonstrate the utility of our proposed methods through simulation studies and an illustrative example assessing effects of antiretroviral therapy on occurrence of AIDS-defining events in patients with HIV from the Vanderbilt Comprehensive Care Clinic.

stat.ME

Equity in the Distribution of Regulatory PM2.5 Monitors

Unequal exposure to air pollution by race and socioeconomic status is well-documented in the U.S. However, there has been relatively little research on inequities in the collection of PM2.5 data, creating a critical gap in understanding which neighborhood exposures are represented in these datasets. In this study we use multilevel models with random intercepts by county and state, stratified by urbanicity to investigate the association between six key environmental justice (EJ) attributes (%AIAN, %Asian %Black, %Hispanic, %White, %Poverty) and proximity to the nearest regulatory monitor at the census tract-level across the contiguous 48 states. We also separately stratify our models by EPA region. Our results show that most EJ attributes exhibit weak or statistically insignificant associations with monitor proximity, except in rural areas where higher poverty levels are significantly linked to greater monitor distances ($β$ = 0.6, 95%CI = [0.49, 0.71]). While the US EPA's siting criteria may be effective in ensuring equitable monitor distribution in some contexts, the low density of monitors in rural areas may impact the accuracy of national-level air pollution monitoring.

stat.AP

Estimating a Causal Exposure Response Function with a Continuous Error-Prone Exposure: A Study of Fine Particulate Matter and All-Cause Mortality

Numerous studies have examined the associations between long-term exposure to fine particulate matter (PM2.5) and adverse health outcomes. Recently, many of these studies have begun to employ high-resolution predicted PM2.5 concentrations, which are subject to measurement error. Previous approaches for exposure measurement error correction have either been applied in non-causal settings or have only considered a categorical exposure. Moreover, most procedures have failed to account for uncertainty induced by error correction when fitting an exposure-response function (ERF). To remedy these deficiencies, we develop a multiple imputation framework that combines regression calibration and Bayesian techniques to estimate a causal ERF. We demonstrate how the output of the measurement error correction steps can be seamlessly integrated into a Bayesian additive regression trees (BART) estimator of the causal ERF. We also demonstrate how locally-weighted smoothing of the posterior samples from BART can be used to create a more accurate ERF estimate. Our proposed approach also properly propagates the exposure measurement error uncertainty to yield accurate standard error estimates. We assess the robustness of our proposed approach in an extensive simulation study. We then apply our methodology to estimate the effects of PM2.5 on all-cause mortality among Medicare enrollees in New England from 2000-2012.

stat.ME

A Calibration Approach to Transportability and Data-Fusion with Observational Data

Two important considerations in clinical research studies are proper evaluations of internal and external validity. While randomized clinical trials can overcome several threats to internal validity, they may be prone to poor external validity. Conversely, large prospective observational studies sampled from a broadly generalizable population may be externally valid, yet susceptible to threats to internal validity, particularly confounding. Thus, methods that address confounding and enhance transportability of study results across populations are essential for internally and externally valid causal inference, respectively. These issues persist for another problem closely related to transportability known as data-fusion. We develop a calibration method to generate balancing weights that address confounding and sampling bias, thereby enabling valid estimation of the target population average treatment effect. We compare the calibration approach to two additional doubly-robust methods that estimate the effect of an intervention on an outcome within a second, possibly unrelated target population. The proposed methodologies can be extended to resolve data-fusion problems that seek to evaluate the effects of an intervention using data from two related studies sampled from different populations. A simulation study is conducted to demonstrate the advantages and similarities of the different techniques. We also test the performance of the calibration approach in a motivating real data example comparing whether the effect of biguanides versus sulfonylureas - the two most common oral diabetes medication classes for initial treatment - on all-cause mortality described in a historical cohort applies to a contemporary cohort of US Veterans with diabetes.

stat.ME

Transporting Experimental Results with Entropy Balancing

We show how entropy balancing can be used for transporting experimental treatment effects from a trial population onto a target population. This method is doubly-robust in the sense that if either the outcome model or the probability of trial participation is correctly specified, then the estimate of the target population average treatment effect is consistent. Furthermore, we only require the sample moments of the effect modifiers drawn from the target population to consistently estimate the target population average treatment effect. We compared the finite-sample performance of entropy balancing with several alternative methods for transporting treatment effects between populations. Entropy balancing techniques are efficient and robust to violations of model misspecification. We also examine the results of our proposed method in an applied analysis of the Action to Control Cardiovascular Risk in Diabetes Blood Pressure (ACCORD-BP) trial transported to a sample of US adults with diabetes taken from the National Health and Nutrition Examination Survey (NHANES) cohort.

stat.ME

A Framework for Covariate Balance using Bregman Distances

A common goal in observational research is to estimate marginal causal effects in the presence of confounding variables. One solution to this problem is to use the covariate distribution to weight the outcomes such that the data appear randomized. The propensity score is a natural quantity that arises in this setting. Propensity score weights have desirable asymptotic properties, but they often fail to adequately balance covariate data in finite samples. Empirical covariate balancing methods pose as an appealing alternative by exactly balancing the sample moments of the covariate distribution. With this objective in mind, we propose a framework for estimating balancing weights by solving a constrained convex program where the criterion function to be optimized is a Bregman distance. We then show that the different distances in this class render identical weights to those of other covariate balancing methods. A series of numerical studies are presented to demonstrate these similarities.

stat.ME